← Back to results

AI Researcher (Multimodal Audio/Video Generation)

Lead research in audio-visual avatar generation at Tavus, shaping the future of human-AI interaction.

Location
San Francisco
Compensation
Not disclosed
Level
senior
Type
full time · Hybrid

Posted by employer 3 months ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Tavus → Save job Scanned from tavus.io

Requirements

Experience
2–3 years
Education
PhD

Not disclosed in this posting: compensation, visa sponsorship.

Joblaze summary

In this role, the Senior AI Researcher at Tavus focuses on advancing audio-visual avatar generation, particularly in conversational contexts. The position requires expertise in generative models, especially diffusion techniques, and proficiency in programming with PyTorch. Ideal candidates possess a PhD and several years of experience in multimodal generation, along with a strong publication record. Tavus, a Series A company, emphasizes innovation in human-computer interaction, making this role suitable for those eager to shape the future of AI.

Joblaze insights

Quick facts

Is the AI Researcher (Multimodal Audio/Video Generation) role remote?
It's hybrid — Tavus expects some on-site time in San Francisco.
How much experience is required?
2–3 years of relevant experience for this AI Researcher (Multimodal Audio/Video Generation) role.
Where is the role based?
Tavus is hiring for this position in San Francisco.
What's the tech stack?
Joblaze extracted these technologies from the posting: Diffusion Models, PyTorch, audio generation, multimodal generation, video generation.
What seniority level is this role?
Tavus targets senior candidates for this position.
Is this full-time or contract?
Full-time for this AI Researcher (Multimodal Audio/Video Generation) role at Tavus.

From the original posting

About Us

Tavus is a research lab pioneering human computing. We’re building AI Humans: a new interface that closes the gap between people and machines, free from the friction of today’s systems. Our real-time human simulation models let machines see, hear, respond, and even look real—enabling meaningful, face-to-face conversations. AI Humans combine the emotional intelligence of humans with the reach and reliability of machines, making them capable, trusted agents available 24/7, in every language, on our terms.

Imagine a therapist anyone can afford. A personal trainer that adapts to your schedule. A fleet of medical assistants that can give every patient the attention they need. With Tavus, individuals, enterprises, and developers can all build AI Humans to connect, understand, and act with empathy at scale.

We’re a Series A company backed by world-class investors including Sequoia Capital, Y Combinator, and Scale Venture Partners.

Be part of shaping a future where humans and machines truly understand each other.

The Role


We’re hiring a Senior AI Researcher to lead research in audio-visual avatar generation. This role is for someone who thrives in ambiguity, has a track record of pushing generative models to new frontiers, and wants to define what human–AI interaction looks like in practice.

Your Mission 🚀

  • Lead research efforts on audio-visual generation for avatars (Neural Avatars, Talking-Heads), with a focus on conversational settings.

  • Design models that are coupled with conversation flow — capturing and generating verbal + non-verbal signals in sync.

  • Drive innovation in diffusion models, long-video generation, and audio-visual modeling.

  • Translate research into production by partnering with Applied ML and engineering.

  • Mentor researchers, set research directions, and publish impactful work.

You’ll Bring:

  • A PhD or equivalent research experience, plus 2–3+ years of hands-on experience applying generative models at scale.

  • Expertise in diffusion models and awareness of the latest efficiency techniques.

  • Experience in multimodal generation — spanning video, audio, and language.

  • Proven innovation in long-video generation and/or audio generation.

  • Excellent programming skills — fluent in PyTorch and GPU-optimized workflows.

  • Track record of publications in top-tier venues (CVPR, NeurIPS, BMVC, ICASSP, etc.).

  • Experience leading research activities or mentoring teams.

Nice-to-Haves:

  • Skills in 3D graphics, Gaussian splatting, or large-scale training setups.

  • Broad exposure to generative AI models beyond your specialty.

  • Familiarity with software development best practices.

Location:


Preferred: San Francisco (hybrid) or London.

Remote within U.S. or Europe considered for exceptional candidates.

Similar positions

Tavus
Senior Software Engineer (CVI)
Tavus · San Francisco
Cartesia
Applied Researcher, Audio
Cartesia · *HQ - San Francisco, CA
Tavus
Data Engineer /ML Ops
Tavus · Remote