← Back to results

Research Engineer, Data Infrastructure (Language Modeling)

Join Cartesia as a Research Engineer to build scalable data infrastructure for language modeling in a collaborative, in-person team.

Location
San Francisco, CA, United States
Compensation
Not disclosed
Level
mid
Type
full time · On-site

Posted by employer 1 day ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Cartesia → Save job Scanned from cartesia.ai

What you'll build

  • Build and operate data processing infrastructure
  • Design and operate data pipelines
  • Run ablation experiments
  • Partner with research and infrastructure teams
  • Establish data quality standards

Must have

  • Hands-on experience with ML data infrastructure
  • Strong modern engineering execution
  • Familiarity with building datasets for generative models

Nice to have

  • Experience with large-scale data processing
  • Experience with pretraining language models

Requirements

Visa
Sponsorship available

Not disclosed in this posting: compensation, years of experience.

Benefits

Meals & Snacks 401(K) Flexible PTO Commuter Allowance Health Insurance Parental Leave

Joblaze summary

In the role of Research Engineer for Data Infrastructure at Cartesia, the individual will focus on developing and maintaining robust data processing systems that support the pretraining of AI models. Key skills include hands-on experience with machine learning data infrastructure and proficiency in modern engineering practices, particularly in building scalable data pipelines. This position is well-suited for someone with a strong background in data systems and a familiarity with generative models, ideally at a mid to senior level. Cartesia emphasizes a collaborative environment, where team members work closely to ensure high-quality data standards that directly influence model performance.

Joblaze insights

  • Listed yesterday — first seen on Joblaze September 25, 2026. Last confirmed on Cartesia's careers page September 25, 2026.
  • Machine Learning appears in 11.4% of 458 comparable mid ai/ml roles in United States; Data Infrastructure appears in 0.2% of 458 comparable mid ai/ml roles in United States.

Quick facts

Is the Research Engineer, Data Infrastructure (Language Modeling) role remote?
No — this is an on-site role in San Francisco, CA, United States.
Where is the role based?
Cartesia is hiring for this position in San Francisco, CA, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: Data Infrastructure, Data Pipelines, Machine Learning, data processing, generative models.
Does Cartesia sponsor work visas for this role?
Yes — the posting indicates visa sponsorship is available for the right candidate.
What seniority level is this role?
Cartesia targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Research Engineer, Data Infrastructure (Language Modeling) role at Cartesia.

From the original posting

About the Role

Data is the lifeblood of our models, and we are looking for a Research Engineer, Data Infrastructure to build the datasets and systems that power pretraining at Cartesia. In this role, you will write performant, scalable infrastructure to acquire, process, and curate massive datasets, and partner closely with research to optimize the characteristics and composition of data mixtures. Your work will directly shape the capabilities and quality of our foundational models.

Your Impact

  • Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets.

  • Design and operate scalable, high-throughput, and reproducible data pipelines — covering ingestion, preprocessing, filtering, deduplication, and augmentation.

  • Design and run ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality.

  • Partner closely with research and infrastructure teams to co-design data loading, versioning, and experimentation pipelines.

  • Establish and enforce rigorous standards for data quality, with a tight feedback loop between dataset characteristics and model behavior.

  • Identify and source novel datasets; manage relationships and budgets with external data vendors and partners.

What You Bring

  • Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training and inference.

  • Strong modern engineering execution: clean, well-tested code, fluency with current tools, and a willingness to pick the right tool for the problem rather than defaulting to familiar patterns.

  • Familiarity with building and evaluating datasets for generative models and reasonable working knowledge of how they're trained and inference.

Nice-To-Haves

  • Experience with large-scale data processing using parallel infrastructure such as Ray, Spark, or Kubernetes.

  • Experience with pretraining language models.

🏦 401(k)

🦖 Your own personal Yoshi

Standard company text repeated across Cartesia's postings is omitted here.

Similar positions

Cartesia
Software Engineer, Data Infrastructure
Cartesia · *HQ - San Francisco, CA
Cartesia
Engineering Manager, Data
Cartesia · *HQ - San Francisco, CA
Cartesia
Analytics Engineer
Cartesia · *HQ - San Francisco, CA
Cartesia
Software Engineer, Platform
Cartesia · *HQ - San Francisco, CA
Cartesia
Researcher, Post Training
Cartesia · *HQ - San Francisco, CA