← Back to results

Build trusted data systems to support clinical operations and analytics at a tech-driven pharma company.

Location
New York, NY, United States
Compensation
$185.5k–$232k/yr
Level
senior
Type
full time · Hybrid

Posted by employer 2 days ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Formation Bio → Save job Scanned from formation.bio

What you'll build

  • Design and operate production data systems
  • Own shared and canonical data models
  • Partner with Product Engineering on application data models
  • Turn recurring data cleaning into maintainable production pipelines
  • Establish strong data quality practices

Must have

  • 5+ years of relevant data engineering experience
  • Experience with pharmaceutical or regulated data
  • Strong Python and SQL skills
  • Experience with Snowflake
  • Experience with Dagster or equivalent
  • Experience integrating complex source data

Nice to have

  • Experience working within validated computerized systems

Practical constraints

  • Hybrid model requiring 3 days per week in office

AI in the day-to-day

Use AI tools, including LLMs and agentic coding systems, to accelerate pipeline development and data quality investigation.

Requirements

Experience
5+ years

Not disclosed in this posting: visa sponsorship.

Benefits

Equity/Stock Options Remote Work Health Insurance

Joblaze summary

In the role of Senior Data Engineer at Formation Bio, the individual will focus on developing and maintaining robust data systems that support various aspects of clinical operations and analytics. Key skills include proficiency in Python and SQL, experience with data modeling and orchestration tools like Dagster, and a solid understanding of regulated data in the biotech sector. This position is ideal for someone with over five years of relevant experience who can effectively collaborate across technical and non-technical teams to enhance data accessibility and quality.

Joblaze insights

  • Listed yesterday — first seen on Joblaze September 19, 2026. Last confirmed on Formation Bio's careers page September 19, 2026.
  • Salary band is in line with the typical range for Data Engineering roles (median ~$180,656).
  • Starts above 82% of 140 comparable senior data engineering roles in United States that list Python we track (median $180,656 across 37 companies). See Python salary trends
  • Python appears in 82.6% of 213 comparable senior data engineering roles in United States; GitHub appears in 1.4% of 213 comparable senior data engineering roles in United States.

Quick facts

Is the Senior Data Engineer role remote?
It's hybrid — Formation Bio expects some on-site time in New York, NY, United States.
What's the salary range?
Formation Bio lists $185,500–$232,000 for this role.
How much experience is required?
At least 5 years of relevant experience for this Senior Data Engineer role.
Where is the role based?
Formation Bio is hiring for this position in New York, NY, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, Dagster, Docker, GitHub, Python, SQL.
What seniority level is this role?
Formation Bio targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Senior Data Engineer role at Formation Bio.

From the original posting

About Formation Bio

About the Position

As a Senior Data Engineer at Formation Bio, you will build the trusted data systems that support clinical operations, drug asset evaluation, business development, analytics, and machine learning through AI Enabled Employees and Agents. You will work across clinical, operational, and third-party data sources to design and operate reliable ingestion pipelines, transformations, data models, and data products.

This role sits at the intersection of Product Engineering, Data Engineering, and Data Science. You will help shape how application data is modeled and exposed, own shared data models and production data products, and prioritize data platform work against clinical and business needs. You will work closely with Product Engineering on application data models and contracts, with Data Science on training datasets and ML use cases, and with human and AI consumers of the data platform.

A data product is not complete merely because it is technically correct or available in a warehouse. It should be understandable to people, usable by applications, useful to Data Science, and structured so AI Enabled Employees and Agents can access it reliably and safely.

Responsibilities

  • Design and operate production data systems that ingest clinical, operational, and third-party vendor data into reliable, queryable data products.
  • Own shared and canonical data models, data contracts, transformations, orchestration, warehouse models, and downstream interfaces.
  • Partner with Product Engineering on application data models, source-system contracts, APIs, events, and data access patterns.
  • Partner with Data Science on productionized training datasets, feature pipelines, data interfaces, and ML use cases.
  • Turn recurring data cleaning, normalization, and transformation work into versioned, tested, observable, and maintainable production pipelines.
  • Build data products for clinical operations, asset evaluation, Business Development, analytics, machine learning, and AI Enabled Employees and Agents.
  • Design data products that are semantically clear, discoverable, machine-readable, permission-aware, traceable, and safe to query.
  • Establish strong data quality, testing, freshness, completeness, lineage, documentation, and observability practices.
  • Own data governance practices for sensitive and regulated data, including access controls, auditability, traceability, and appropriate data handling.
  • Participate in support and incident response for data platform issues, including diagnosis, stakeholder communication, remediation, and prevention of recurrence.
  • Use AI tools, including LLMs and agentic coding systems, to accelerate pipeline development, data modeling, debugging, documentation, and data quality investigation while validating their output.
  • Contribute to architecture and design reviews, mentor other engineers, and improve the engineering practices used across the organization.

About You

  • 5+ years of relevant data engineering experience building and operating production data systems.
  • Experience with pharmaceutical, biology, HIPPA or other regulated data core to BioTech is required.
  • Strong Python and SQL skills, with deep experience in data modeling and warehouse systems, especially Snowflake.
  • Experience with Dagster as an orchestration systems (or equivalent) transformation tooling such as dbt or an equivalent approach.
  • Experience with data contracts, schema evolution, data quality testing, observability, lineage, and production incident response.
  • Experience integrating messy clinical, operational, vendor, or otherwise complex source data.
  • Working knowledge of Docker, GitHub, and Terraform or OpenTofu sufficient to partner effectively with SRE.
  • Experience building data products and access patterns for applications, Data Science, analytics, human users, and AI Enabled Employees and Agents.
  • Strong judgment about when to build reusable platform capabilities versus one-off stakeholder solutions.
  • Daily fluency with AI tools and the ability to validate generated code, transformations, and data-modeling decisions.
  • Exceptional collaboration and communication skills across Product Engineering, Data Science, Clinical Operations, Data Management, Business Development, and other non-technical partners.
  • Experience working within and building validated computerized systems (CSV) is a plus.

Total Compensation Range: $185,500 - $232,000

Compensation Individual compensation is determined by several factors, including role scope, geographic location, and skills & experience. Your offer will reflect where you fall within the range based on these considerations. In addition to base salary, we offer equity, comprehensive benefits, and generous perks. If the posted range doesn't match your expectations, we still encourage you to apply!

Standard company text repeated across Formation Bio's postings is omitted here.

Similar positions

Formation Bio
Senior Data Scientist
Formation Bio · New York, NY; Boston, MA
Formation Bio
Senior Software Engineer
Formation Bio · New York, NY, United States
Formation Bio
Associate Director, Data Engineering
Formation Bio · New York, NY; Boston, MA; San Francisco, CA
Formation Bio
Senior Site Reliability Engineer
Formation Bio · New York, NY, United States
Formation Bio
Senior Data Scientist - Real World Data
Formation Bio · New York, NY; Boston, MA; San Francisco, CA; Raleigh-Durham, NC