← Back to results

Machine Learning Engineer (Real-Time Speech Translation)

Join Lilt as a Machine Learning Engineer to build a real-time speech translation backend for a cutting-edge live translation product.

Location
Washington, D.C., United States
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 15 hours ago

First seen on Joblaze 3 hours ago

Last verified on the company career page 3 hours ago

Apply at Lilt → Save job Scanned from lilt.com

What you'll build

  • Build and manage services for real-time audio and text streaming
  • Integrate and serve streaming speech recognition and machine translation models
  • Develop logic for model-based confidence scoring
  • Architect and scale production ML infrastructure on Kubernetes
  • Establish comprehensive instrumentation for real-time performance

Must have

  • 3+ years building production backend or ML serving systems in Python
  • Strong async programming skills
  • Hands-on experience with real-time streaming transport
  • Experience serving ML models in production on GPUs
  • Experience integrating speech or NLP models into production systems

Nice to have

  • Ray Serve specifically
  • Familiarity with simultaneous or incremental MT concepts
  • Machine translation quality estimation
  • Message brokers for real-time fan-out
  • Streaming text-to-speech integration

Practical constraints

  • US citizenship and residence in the United States

Role intensity

70% hands-on coding

AI in the day-to-day

We combine AI-driven PR reviews to accelerate development.

Requirements

Experience
3+ years
Education
Bachelor's degree
Visa
No sponsorship (stated in posting)

Not disclosed in this posting: compensation, work arrangement.

Joblaze summary

In this role, the Machine Learning Engineer will develop and manage the backend for a new real-time speech translation product, ensuring seamless integration of audio input and translated output. Key skills include expertise in Python, real-time streaming technologies, and experience with ML model deployment on GPU-accelerated Kubernetes clusters. This position is ideal for someone with a strong background in backend engineering and a focus on low-latency systems, particularly those who thrive in collaborative environments with a clear product vision.

Joblaze insights

  • Listed today — first seen on Joblaze October 3, 2026. Last confirmed on Lilt's careers page October 3, 2026.
  • Python appears in 48.3% of 464 comparable mid ai/ml roles in United States; Machine Translation appears in 0.2% of 464 comparable mid ai/ml roles in United States.

Quick facts

How much experience is required?
At least 3 years of relevant experience for this Machine Learning Engineer (Real-Time Speech Translation) role.
What's the tech stack?
Joblaze extracted these technologies from the posting: ASR, Kubernetes, Machine Translation, NLP, Python, Ray Serve.
What seniority level is this role?
Lilt targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Machine Learning Engineer (Real-Time Speech Translation) role at Lilt.

From the original posting

Role Summary

We are building a new live translation product. We are looking for an ML Engineer to build the real-time speech translation backend that powers it.

You will own the real-time speech translation backend end-to-end, from live audio input to translated output. You will build on LILT's production model serving platform (Ray Serve on GPU Kubernetes clusters) and our in-house adaptive machine translation models, working closely with the senior architects of that platform and with our language processing researchers. The ASR and MT models exist. Your job is to make them work together as a low-latency streaming system that holds up in production.

This is a hands-on backend engineering role for those looking to own a real-time ML system from end to end, supported by expert guidance and a clear product vision. Our team adopts an AI-first approach, leveraging agentic coding and AI-driven PR reviews to accelerate development. We combine this with deep technical expertise, requiring not only expert Python proficiency but also a comprehensive understanding of the entire ML stack, from optimizing neural network architectures to managing production infrastructure on Kubernetes, and making informed, cost-aware decisions on hardware selection.

Location & eligibility: This position requires US citizenship and residence in the United States. Preferred locations are Washington, D.C.; Boston, MA; and Indianapolis, IN (East Coast / ET timezone preferred).

Key Responsibilities

  • Real-time pipeline architecture: Build and manage services for high-throughput, real-time audio and text streaming. Handle signal processing, session lifecycles, and concurrency management to ensure robust operation under load.

  • ML model integration: Integrate and serve streaming speech recognition and machine translation models, collaborating with research teams to ensure models operate within required latency budgets.

  • Quality and confidence workflows: Develop logic for model-based confidence scoring, routing segments for human intervention as needed, and broadcasting real-time updates and corrections to end-users.

  • Infrastructure and scale: Architect and scale production ML infrastructure on GPU-accelerated Kubernetes clusters. Implement batching, load balancing, and autoscaling strategies to maintain performance and cost-efficiency.

  • Latency engineering: Establish comprehensive instrumentation for real-time performance. Identify bottlenecks, optimize system throughput, and drive down end-to-end latency metrics to meet production standards.

  • Interface and API definition: Define technical contracts and interfaces for audio ingestion and downstream service integrations. Partner with frontend and platform engineering teams to maintain clean, robust integration points.

  • Collaboration and technical leadership: Drive cross-team alignment by defining clear API interfaces and technical contracts, facilitating effective communication between engineering and product teams to ensure seamless system integration.

Required Qualifications

  • BS or MS in Computer Science or a related field, or equivalent practical experience.

  • 3+ years building production backend or ML serving systems in Python, including strong async programming (asyncio) skills.

  • Hands-on experience with real-time streaming transport: WebSocket or gRPC bidirectional streaming, session state, backpressure, and connection lifecycle handling.

  • Experience serving ML models in production on GPUs (Ray Serve, Triton, vLLM, or similar), with Docker and Kubernetes.

  • Experience integrating speech or NLP models into production systems, ideally streaming ASR (partial hypotheses, endpointing, VAD).

  • A latency-engineering mindset: you have profiled, instrumented, and optimized a real-time or low-latency system and can reason in per-stage budgets.

  • Effective use of AI coding agents (Claude Code, Codex, or similar) on top of fundamentals learned the hard way: you let agents do the typing, but you can debug, review, and reason about every line without them, and you know when not to trust them.

  • US citizenship and residence in the United States (contract requirement).

Preferred Qualifications

  • Ray Serve specifically, including streaming responses and model multiplexing.

  • Familiarity with simultaneous or incremental MT concepts (retranslation, prefix stability, wait-k policies).

  • Machine translation quality estimation (COMET/CometKiwi class models) or other confidence estimation in production.

  • Message brokers for real-time fan-out and state distribution (RabbitMQ or similar).

  • Streaming text-to-speech integration and time-to-first-audio optimization.

  • WebRTC and SFU concepts, or voice pipeline frameworks (LiveKit Agents, Pipecat).

  • Handling of CJK and other non-Latin text in NLP pipelines (our first languages are Japanese, Korean, and English).

  • Observability tooling (Datadog, Prometheus) for production ML systems.

Our Tech

What sets our platform apart:

  • Brand-aware AI that learns your voice, tone, and terminology to ensure every translation is accurate and consistent


LILT in the News

Standard company text repeated across Lilt's postings is omitted here.

Similar positions

Lilt
Backend Engineer
Lilt · Washington, D.C., United States
Lilt
Senior Full Stack Engineer
Lilt · Washington, D.C., United States
Lilt
Senior DevOps Engineer
Lilt · Washington D.C.