← Back to results

Data Infrastructure Engineer

Join Mind Robotics as a Data Infrastructure Engineer to build and scale data pipelines for high-dimensional sensor data.

Location
Palo Alto
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 7 months ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Mind Robotics → Save job Scanned from mindrobotics.com

Requirements

Experience
2+ years

Not disclosed in this posting: compensation, work arrangement, visa sponsorship.

Joblaze summary

In the role of Data Infrastructure Engineer at Mind Robotics, the individual will focus on developing and maintaining data pipelines that transform raw sensor data into usable training datasets for robotic systems. Key skills include proficiency in Python and experience with distributed data processing frameworks, as well as a solid understanding of data storage concepts. This position is ideal for someone with a couple of years in software engineering, particularly in data engineering, who thrives in collaborative environments and is eager to take ownership of complex systems. Mind Robotics emphasizes hands-on problem-solving in robotics, making this a unique opportunity for those passionat

Joblaze insights

Quick facts

How much experience is required?
At least 2 years of relevant experience for this Data Infrastructure Engineer role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Airflow, Dagster, Dask, Flink, Prefect, Python.
What seniority level is this role?
Mind Robotics targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Data Infrastructure Engineer role at Mind Robotics.

From the original posting

About Mind:

Mind Robotics is building Physical AI for real-world industrial deployment, starting with the factory floor. We believe the hardest problems in AI are solved when researchers and engineers are hands-on with the physical world every day - and we're looking for people who are passionate about robotics, value ownership, and are excited to tackle difficult problems. Join us if you want to move beyond digital intelligence and put intelligence into motion.

About the team and the role:

Mind Robotics is building robots that learn from real-world experience. That starts with a data engine: the pipelines and infrastructure that turn raw, messy, multimodal sensor streams from robots and human demonstrations into high-quality, well-curated training data at scale.

As a Data Infrastructure Engineer, you'll build and operate the pipelines that turn raw sensor and demonstration data into training-ready datasets — from ingestion off real robots and capture devices, through processing and quality filtering, to the dataloaders that feed model training. The systems work today; your job is to build within the architecture, take ownership of specific pipelines and services, and help harden the system as it scales.

Responsibilities:

  • Build and scale ingestion pipelines for high-volume, high-dimensional sensor data.

  • Build and maintain data quality and curation systems, including both manual review workflows and auto-labeling.

  • Contribute to storage and retrieval design for large-scale multimodal datasets, including format choices and versioning/lineage.

  • Build and operate distributed processing infrastructure (batch and streaming).

  • Debug production issues in live data pipelines — data corruption, schema drift, backpressure, and failures that only show up at scale.

  • Work with modeling/research partners to understand data quality, format, and structure needs, and translate them into working pipelines.

  • Participate in design and code review across the data infrastructure stack.

Requirements:

  • 2+ years of software engineering experience, with some exposure to data pipelines, data engineering, or backend systems.

  • Strong programming fundamentals in Python, with the ability to write performant, production-grade data processing code.

  • Experience with at least one distributed data processing framework (e.g., Spark, Ray, Dask, or Flink), or strong fundamentals and willingness to ramp up quickly.

  • Familiarity with data storage concepts — object storage, data lake table formats, and warehouse vs. lake tradeoffs.

  • Bias for ownership: you've taken features or systems from prototype to production.

  • Clear communicator who collaborates well with research/modeling partners and more senior teammates.

  • Experience with workflow orchestration tools (e.g., Airflow, Dagster, Prefect) is a plus.

  • Experience with streaming ingestion for high-volume, near-real-time data is a plus.

  • Experience building automated data quality/curation systems (statistical filtering, anomaly detection, deduplication, or using ML models in an annotation/filtering pipeline) is a plus.

  • Experience with robotics- or embodied-AI-specific data (multi-embodiment datasets, teleoperation/demonstration data, egocentric video) is a plus.

Similar positions

Mind Robotics
Research Engineer
Mind Robotics · Palo Alto
Mind Robotics
DevOps Engineer
Mind Robotics · Palo Alto
Mind Robotics
Robotics Software Engineer
Mind Robotics · Palo Alto
Mind Robotics
Application Engineer
Mind Robotics · Palo Alto
Mind Robotics
Research + Modeling
Mind Robotics · Palo Alto