← Back to results

Research Scientist, Video Foundation Models

Join Cantina as a Research Scientist to develop next-generation video foundation models and shape the future of AI-driven creativity.

Location
California, United States
Compensation
$200k–$320k/yr
Level
senior
Type
full time

Posted by employer 2 months ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Cantina → Save job Scanned from cantina.com

What you'll build

  • Research and develop native video and multimodal foundation models
  • Explore new model architectures and training objectives
  • Build and improve large-scale data curation and training pipelines
  • Design systematic experiments for model evaluation
  • Collaborate with teams to shape technical roadmap

Must have

  • Strong research and engineering experience in generative modeling
  • Hands-on experience training large-scale models
  • Track record of developing impactful models or systems
  • Ability to independently own ambiguous research problems
  • Specialized depth in foundation model lifecycle areas

Not disclosed in this posting: years of experience, work arrangement, visa sponsorship.

Benefits

401k Match Lifestyle Spending Account Equity/Stock Options Paid Time Off Health Insurance Parental Leave

Joblaze summary

In this role, the Research Scientist will focus on developing and scaling native video and multimodal foundation models, engaging in all stages from architecture design to post-training evaluation. Key skills include expertise in generative modeling and hands-on experience with large-scale training using deep learning frameworks. This position is well-suited for individuals with a strong research background and the ability to tackle complex problems independently while collaborating with a diverse team. Cantina is in an early stage of building its core team, emphasizing innovation in social AI.

Joblaze insights

Quick facts

What's the salary range?
Cantina lists $200,000–$320,000 for this role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Deep Learning, Multimodal Models, generative modeling, video generation.
What seniority level is this role?
Cantina targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Research Scientist, Video Foundation Models role at Cantina.

From the original posting

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation.

In this role, you will work on foundational research and large-scale model development across the full model lifecycle, including architecture, data, training, evaluation, post-training, training systems, and efficient inference. You will have the opportunity to shape both the technical direction and the team from an early stage.

What You’ll Do:

  • Research, develop, and scale native video and multimodal foundation models, from early prototypes through large-scale pre-training, continued training, and post-training.

  • Explore new model architectures, training objectives, and conditioning mechanisms for video generation, reference- and memory-based generation, multimodal interaction, and joint audio-video generation.

  • Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines for high-quality and controllable generation.

  • Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency.

  • Collaborate closely with researchers, engineers, and product teams to help shape the technical roadmap and, where appropriate, translate model advances into real-world capabilities.

  • Contribute to research publications and open-source releases when appropriate.

What You’ll Bring:

  • Strong research and engineering experience in generative modeling, including areas such as diffusion models, flow matching, DiTs, video generation, multimodal models, world models, or related fields.

  • Hands-on experience training and evaluating large-scale image, video, or unified multimodal models using modern deep learning frameworks and distributed training systems.

  • A strong track record of developing impactful models or systems, demonstrated through research publications, open-source contributions, production impact, or other significant technical work.

  • Ability to independently own ambiguous research problems, move effectively from ideas to experiments, and work well in a highly collaborative environment.

  • Specialized depth in one or more areas across the foundation model lifecycle, such as model architecture, data curation, controllable generation, multimodal understanding and conditioning, post-training and reward modeling, model acceleration, inference systems, or deployment.

Compensation:

The anticipated annual base salary range for this role is between $200,000-$320,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Similar positions

Cantina
Cantina
Cantina
Cantina
Research Intern
Cantina · Singapore