← Back to results

Machine Learning Engineer, Ops

Join Cantina as an MLOps Engineer to build and scale inference infrastructure for generative audio models.

Location
Remote (U.S. or Europe)
Compensation
$125k–$165k/yr
Level
mid
Type
full time

Posted by employer 1 month ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Cantina → Save job Scanned from cantina.com

Not disclosed in this posting: years of experience, work arrangement, visa sponsorship.

Benefits

401k Match Equity/Stock Options Paid Time Off Health Insurance Parental Leave

Joblaze summary

The Machine Learning Engineer in Operations at Cantina focuses on building and scaling the inference infrastructure for generative audio models, ensuring efficient model serving for both streaming and batch applications. Key skills include expertise in Kubernetes for orchestration, MLOps practices, and a solid understanding of audio model architectures like TTS and ASR. This role is ideal for someone with a strong technical background in software engineering and experience in high-performance systems. Cantina's innovative environment emphasizes collaboration between research and production teams to enhance audio model performance.

Joblaze insights

Quick facts

What's the salary range?
Cantina lists $125,000–$165,000 for this role.
What's the tech stack?
Joblaze extracted these technologies from the posting: CI/CD, GPU, Go, Kubernetes, MLOps, Python.
What seniority level is this role?
Cantina targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Machine Learning Engineer, Ops role at Cantina.

From the original posting

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

 

About the Role:

We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.

What You’ll Do:

  • Design and maintain inference infrastructure for generative audio model architectures.

  • Implement and manage high-performance inference engines.

  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.

  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.

  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.

  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.

  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring:

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.

  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.

  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.

  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.

  • Experience with GPU-accelerated inference and performance profiling techniques.

  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation:

The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

 

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Similar positions

Cantina
Cantina
Machine Learning Engineer - Voice Conversion
Cantina · Remote (U.S. or Europe)
Cantina
Cantina
Cantina