Join Cartesia as a Research Engineer to build scalable data infrastructure for language modeling in a collaborative, in-person team.
Posted by employer 1 day ago
First seen on Joblaze 1 day ago
Last verified on the company career page 1 day ago
Skills & Technologies
What you'll build
Must have
Nice to have
Requirements
Not disclosed in this posting: compensation, years of experience.
Benefits
Joblaze summary
In the role of Research Engineer for Data Infrastructure at Cartesia, the individual will focus on developing and maintaining robust data processing systems that support the pretraining of AI models. Key skills include hands-on experience with machine learning data infrastructure and proficiency in modern engineering practices, particularly in building scalable data pipelines. This position is well-suited for someone with a strong background in data systems and a familiarity with generative models, ideally at a mid to senior level. Cartesia emphasizes a collaborative environment, where team members work closely to ensure high-quality data standards that directly influence model performance.
Joblaze insights
Quick facts
From the original posting
Data is the lifeblood of our models, and we are looking for a Research Engineer, Data Infrastructure to build the datasets and systems that power pretraining at Cartesia. In this role, you will write performant, scalable infrastructure to acquire, process, and curate massive datasets, and partner closely with research to optimize the characteristics and composition of data mixtures. Your work will directly shape the capabilities and quality of our foundational models.
Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets.
Design and operate scalable, high-throughput, and reproducible data pipelines — covering ingestion, preprocessing, filtering, deduplication, and augmentation.
Design and run ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality.
Partner closely with research and infrastructure teams to co-design data loading, versioning, and experimentation pipelines.
Establish and enforce rigorous standards for data quality, with a tight feedback loop between dataset characteristics and model behavior.
Identify and source novel datasets; manage relationships and budgets with external data vendors and partners.
Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training and inference.
Strong modern engineering execution: clean, well-tested code, fluency with current tools, and a willingness to pick the right tool for the problem rather than defaulting to familiar patterns.
Familiarity with building and evaluating datasets for generative models and reasonable working knowledge of how they're trained and inference.
Experience with large-scale data processing using parallel infrastructure such as Ray, Spark, or Kubernetes.
Experience with pretraining language models.
🏦 401(k)
🦖 Your own personal Yoshi
Standard company text repeated across Cartesia's postings is omitted here.