← Back to results

Machine Learning Engineer

Join Hippocratic AI as a Machine Learning Engineer to build a self-improvement system focused on safety in healthcare.

Location
Menlo Park, CA, United States
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 2 days ago

First seen on Joblaze 9 hours ago

Last verified on the company career page 9 hours ago

Apply at Hippocratic AI → Save job Scanned from hippocraticai.com

What you'll build

  • Build and maintain training and evaluation loops
  • Design and implement reward and feedback signals
  • Build evaluation harnesses and metrics
  • Own data pipelines and automated data flywheels
  • Debug model-quality regressions

Must have

  • Strong MLE fundamentals
  • Excellent Python and clean ML training code
  • Solid grasp of data pipelines and large-scale training
  • Hands-on experience with a feedback or learning loop
  • Ran a retraining or continual-learning pipeline
  • Reinforcement learning foundations

Nice to have

  • PhD or MS in RL / ML
  • Experience at a lab or company doing RLHF
  • Familiarity with LLM fine-tuning
  • Experience with agent orchestration

Requirements

Education
Master's degree

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Joblaze summary

The Machine Learning Engineer at Hippocratic AI focuses on developing and maintaining the core training and evaluation loops for a self-improving machine learning system, ensuring reliability and reproducibility. Key skills include strong foundations in machine learning, proficiency in Python, and experience with feedback loops and data pipelines. This role is suited for candidates with a solid background in reinforcement learning and practical experience in deploying ML systems. The company is dedicated to creating a safety-focused AI platform for healthcare, backed by significant funding and expertise.

Joblaze insights

  • Listed today — first seen on Joblaze September 25, 2026. Last confirmed on Hippocratic AI's careers page September 25, 2026.
  • Python appears in 47.3% of 459 comparable mid ai/ml roles in United States; Data Pipelines appears in 2.4% of 459 comparable mid ai/ml roles in United States.

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: Data Pipelines, Machine Learning, Python, feedback loops, reinforcement learning.
What seniority level is this role?
Hippocratic AI targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Machine Learning Engineer role at Hippocratic AI.

From the original posting

About the role

We are building a recursive self-improvement system — a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives.

This is an engineering-first role with deep reinforcement learning requirements. You should be equally comfortable writing robust production ML code and reasoning about reward design, credit assignment, and why feedback-driven systems become unstable.

What you'll do

  • Build and maintain the training, evaluation, and deployment loops at the core of the self-improvement system, with a strong emphasis on reproducibility and reliability.

  • Design and implement reward and feedback signals; investigate and mitigate reward hacking, specification gaming, and distribution drift.

  • Build evaluation harnesses and metrics before models — because a self-improving system is only as safe as its measurement of “better.”

  • Own data pipelines and automated data flywheels that feed the learning loop.

  • Debug subtle model-quality regressions and stabilize training and feedback loops that go non-stationary.

  • Collaborate with research and product to turn methods into robust, shippable systems.

What we're looking for

Must-Haves:

Strong MLE fundamentals (non-negotiable)

  • Excellent Python and clean, well-tested ML training code.

  • Solid grasp of data pipelines, distributed / large-scale training, and experiment tracking.

  • The instinct and skill to debug why a model silently got worse — not just why it crashed.

Hands-on experience with a feedback or learning loop (at least one).

  • Built or owned part of a feedback loop — a reward model, an evaluation harness, or the data pipeline for an RLHF/RLAIF or active-learning system.

  • Ran a retraining or continual-learning pipeline where a model consumed its own predictions or production data (e.g. ranking, recommendations, fraud, spam).

  • Fine-tuned LLMs with human or AI feedback, or built agentic evaluation harnesses.

Reinforcement learning foundations and curiosity.

  • Working knowledge of reward modeling, on-policy vs. off-policy tradeoffs, and credit assignment (does not need to be a research-level RL expert).

  • Has seen — or can reason clearly about — feedback-system failure modes: reward hacking, specification gaming, feedback loops amplifying errors.

  • Comfortable evaluating non-stationary systems (systems whose behavior and data distribution change over time).

Systems and evaluation instinct.

  • Builds the eval before the model; treats measurement as a first-class deliverable.

  • Has shipped an ML system into production and kept it healthy over time.

Strongest signal

Bonus, not required:

The ideal candidate has built or shipped a full system that improved from its own outputs or feedback end to end. This is rare at this level, so treat it as a standout differentiator rather than a filter. Examples:

  • RLHF / RLAIF pipelines

  • Self-play systems

  • Active-learning loops

  • Automated data flywheels

  • Agentic evaluation harnesses

Nice to Have:

  • PhD or MS in RL / ML paired with real production experience (either the science or the engineering half alone is fine if the other is strong).

  • Experience at a lab or company doing RLHF, agents, or large-scale ML infrastructure.

  • Familiarity with LLM fine-tuning, evaluation frameworks, or agent orchestration.

Standard company text repeated across Hippocratic AI's postings is omitted here.

Similar positions

HappyRobot
Machine Learning Engineer
HappyRobot · San Francisco
Periodic Labs
ML Systems Engineer
Periodic Labs · Menlo Park, CA
Mirendil