This position may no longer be available
This job was last seen 1 week ago. The listing may have been removed by the employer.
Own and evolve the data systems at WorkOS to ensure reliable and scalable data management across the organization.
Posted by employer 1 week ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 week ago
Skills & Technologies
What you'll build
Must have
Nice to have
AI in the day-to-day
Comfortable using LLMs and coding agents as part of your day-to-day development workflow.
Requirements
Not disclosed in this posting: compensation, visa sponsorship.
Benefits
Joblaze summary
In this role, the Data Engineer at WorkOS is responsible for managing and enhancing the internal data platform, focusing on data ingestion, orchestration, and governance. Key skills include expertise in Snowflake, dbt, and data orchestration tools, along with strong SQL and Python capabilities. This position is ideal for experienced engineers who thrive in collaborative environments and can navigate complex data challenges while maintaining high standards of data quality and accessibility. The team operates with a lean structure, emphasizing rapid experimentation and open communication.
Joblaze insights
Quick facts
From the original posting
About WorkOS 🚀
The Data team at WorkOS plays a central role on the company operations to ensure information and insights are accurate and available from any surface, whether that's a visualization tool or an agent's response, to accelerate company growth and scale. The team owns WorkOS's internal data platform end to end: ingestion, orchestration, the Snowflake warehouse, dbt transformations, governance and access controls. We also own the consumption layer: the reverse-ETL syncs that land data in Salesforce and Slack, the semantic views that agents query, the definitions that keep reporting consistent across the company, and the visualization tooling the company uses (ask us about Flashboards).
We run on an agent-first operating system: we are a lean team that moves fast, we document so agents can execute, and AI agents query the warehouse, run our runbooks, and open pull requests alongside us. We trust each other's expertise, we work out loud and in the open, and not opposed to rapid experimentation to get to the right solution for the business.
We're hiring a Data Engineer to own and evolve the systems that move, transform, protect, and serve data across our internal warehouse. This is a high-ownership role on a lean team: you will be the DRI for the platform and its core data models, partner directly with Product Engineering, RevOps, Finance, GTM Engineering, and Security, and raise the bar on reliability, correctness, and operational rigor. The role spans the platform and what runs on it. You will own ingestion, orchestration, and access governance; the dbt models and metric definitions that depend on them; and how AI is applied across the data stack, so that trusted data is easy for both people and AI agents to discover, understand, and work with.
You don't need to have done this exact job before. The best data engineers at WorkOS are the ones who notice in a Slack thread that a number doesn't match, trace it from the dashboard through the Gold model to the connector, fix the connector, and confirm the definition with its owner before the weekly review. When a business team asks for a field in Salesforce, they don't hand it off; they build the model, write the sync, and add the test. They would rather automate the runbook than run it a third time. They enjoy collaboration with different parts of the business and can find simple, scalable solutions to ambiguous problems. If that sounds like how you work, we want to talk.
Own the reliability, freshness, and scaling of the ingestion and orchestration pipelines that land source data in Snowflake, including monitoring, alerting, runbooks, and backfill and reprocessing patterns
Design, build, and scale the core dbt models (Bronze, Silver, Gold) for billing and usage, CRM, product events, and GTM funnel reporting
Partner with Product, Finance, RevOps, and GTM to define metrics and codify them in the warehouse, and diagnose and resolve data quality and freshness issues at source, not just downstream
Own Snowflake RBAC, dynamic masking, and PII classification for humans, agents, and service accounts, so sensitive data is protected without blocking legitimate use, and review DDL and access requests from Engineering and GTM
Own the reverse-ETL layer and the runbooks that deliver warehouse data to Salesforce, Slack, and internal agents
Extend the semantic views, context, and evaluations that let agents answer business questions accurately, and expand what agents can operate directly, from runbooks and ingestion to pull requests
Build and advance the CI/CD review gates in the data-platform monorepo, including the automated review that agent-authored pull requests pass through
Own the infrastructure the data platform runs on: the compute, deployments, secrets and access patterns, and environments behind ingestion and orchestration, with attention to availability and failure modes, in partnership with our infrastructure engineers
Example problems you might work on, possible paths rather than a fixed roadmap:
Standardizing how Postgres and SaaS sources are ingested, and consolidating end-to-end pipeline deployment and orchestration in Prefect
Managing Snowflake RBAC, resource management, and masking policies as code with Terraform
Using query history to improve analytics tables and semantic views, and to keep agent answers accurate and consistent as data grows and metrics shift
Automating masking coverage as new data sources arrive, so adding a new table no longer needs a human in the loop
Automating near-certain matches in the Identity Graph
Pseudonymizing product data in the ingestion path, alongside compliance and data-deletion policies
Extending transcript aggregation across sources, including PII detection before transcripts enter the warehouse
Building the reverse-ETL framework and runbooks that deliver data to other tooling and agents
Running self-hosted data services on Kubernetes and managing their AWS resources (IAM, storage, secrets) as code with Terraform
5+ years (or equivalent) building and operating production data platforms, including the transformation layer
Deep experience with Snowflake, dbt, and an orchestrator such as Prefect, Airflow, or Dagster, and with ingesting data from production databases and SaaS sources, including backfill and reprocessing patterns
Experience running data systems in production: monitoring, alerting, debugging, incident response habits, and pragmatic SLO/SLA thinking
Experience with warehouse access governance: RBAC, masking policies, and PII handling
Strong data modeling judgment and an understanding of schema evolution
Strong SQL and Python; disciplined engineering practices (tests, docs, reviews, CI/CD)
Proven experience operating as the only engineer on a layer, balancing foundational architecture with urgent business needs
Working knowledge of how GTM, Finance, and Product teams operate in addition to working with Engineering, and experience translating ambiguous business questions into technical specs
Comfortable using LLMs and coding agents as part of your day-to-day development workflow, with the judgment to validate their output and scale it into shared processes, and embracing AI and automation to scale platform work and accelerate your teammates
A systems thinker who reasons carefully about freshness, correctness, and failure modes, especially for data that everything else depends on
Pragmatic: you start with the simplest solution that works, prove it, then scale it, and you balance fast answers with durable solutions
You look for where AI can remove a bottleneck for the people and agents who depend on the warehouse, and you build the durable tools and workflows they rely on every day
Benefits and Perks (US Only) 💖
401k matching
Competitive Equity
Healthcare, dental and vision coverage
FSA, ST/LT Disability, Voluntary Life
Carrot fertility benefits
12 weeks fully paid parental leave
Commuter benefits for hybrid employees in SF/NYC
Unlimited token usage!
Standard company text repeated across WorkOS's postings is omitted here.