← Back to results

Multimodal AI Model Optimization Research Engineer

Join Tavus as a Research Engineer to optimize cutting-edge multimodal AI models for production readiness.

Location
Remote
Compensation
Not disclosed
Level
senior
Type
full time · Hybrid

Posted by employer 3 days ago

First seen on Joblaze 2 days ago

Last verified on the company career page 12 hours ago

Apply at Tavus → Save job Scanned from tavus.io

What you'll build

  • Optimize research models for production
  • Define metrics and run experiments
  • Benchmark trade-offs across latency, cost, and quality
  • Collaborate with researchers and engineers

Must have

  • Strong experience in deep learning using PyTorch
  • Hands-on experience with model optimization and compression
  • Strong understanding of inference performance and GPU fundamentals
  • Strong Python coding skills

Nice to have

  • Optimization of diffusion models
  • Experience with real-time or streaming systems
  • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
  • Experience writing custom Triton/CUDA kernels

AI in the day-to-day

Our models power everything from text-to-video AI avatars to real-time conversational video experiences.

Not disclosed in this posting: compensation, years of experience, visa sponsorship.

Benefits

Flexible work schedules Gear Stipends Unlimited PTO Competitive Healthcare

Joblaze summary

The Multimodal AI Model Optimization Research Engineer at Tavus focuses on enhancing the efficiency and readiness of advanced AI models for production through techniques like sparsification and quantization. Key skills include deep learning expertise in PyTorch and hands-on experience with model optimization and compression methods. This role is suited for experienced professionals who thrive in fast-paced startup environments and possess strong collaboration and communication abilities. Tavus emphasizes a culture of diversity and innovation, aiming to redefine human-AI interaction.

Joblaze insights

  • Listed 2 days ago — first seen on Joblaze September 29, 2026. Last confirmed on Tavus's careers page October 1, 2026.
  • PyTorch appears in 18.3% of 536 comparable senior ai/ml roles in United States; quantization appears in 0.2% of 536 comparable senior ai/ml roles in United States.

Quick facts

Is the Multimodal AI Model Optimization Research Engineer role remote?
It's hybrid — Tavus expects some on-site time in Remote.
Where is the role based?
Tavus is hiring for this position in Remote.
What's the tech stack?
Joblaze extracted these technologies from the posting: CUDA, Deep Learning, GPU, Pruning, PyTorch, Triton.
What seniority level is this role?
Tavus targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Multimodal AI Model Optimization Research Engineer role at Tavus.

From the original posting

About Us

At Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable.

We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.

By enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.

We are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.

The Role

We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team.

Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.

Your Mission

  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization

  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality

  • Partner closely with researchers and engineers to turn new ideas into deployable systems

Requirements

  • Strong experience in deep learning using PyTorch

  • Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision

  • Understanding of efficient architectures such as low-rank adapters

  • Strong understanding of inference performance and GPU/accelerator fundamentals

  • Strong Python coding skills and reliable research engineering practices

  • Experience working with large models and datasets in cloud environments

  • Ability to read ML papers, reproduce results, and adapt ideas

  • Clear communication and collaboration skills

Preferred Experience

  • Optimization of diffusion models, video/audio generative models, or large language models

  • Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)

  • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA

  • Experience writing custom Triton/CUDA kernels or low-level performance tuning

  • Experience with experiment tracking, benchmarking, and profiling at scale

  • Prior experience in research engineering or applied science roles

Location

This position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.

Benefits

When you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.

Culture & Diversity

We are not looking for cultural fits — we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients..

Similar positions