← Back to results

Research Scientist, Foundation Model (Video Generation)

Join Pika as a lead Research Scientist to advance real-time multimodal foundation models for creative technology.

Location
Palo Alto HQ
Compensation
Not disclosed
Level
lead
Type
full time · Hybrid

Posted by employer 3 months ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Pika → Save job Scanned from pika.art

Skills & Technologies

Requirements

Experience
5+ years
Education
PhD

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

401k Match Equity/Stock Options Health Insurance

Joblaze summary

In this role, the Research Scientist will focus on leading the development of large-scale multimodal foundation models, specifically in pre-training and mid-training processes. The position requires expertise in generative architectures and hands-on experience with programming tools like Python and PyTorch, as well as a strong background in dataset curation. This opportunity is ideal for seasoned researchers with a proven publication record and a passion for advancing creative technology. Pika's collaborative environment emphasizes innovation and aims to make generative technology accessible to a wider audience.

Joblaze insights

Quick facts

Is the Research Scientist, Foundation Model (Video Generation) role remote?
It's hybrid — Pika expects some on-site time in Palo Alto HQ.
How much experience is required?
At least 5 years of relevant experience for this Research Scientist, Foundation Model (Video Generation) role.
Where is the role based?
Pika is hiring for this position in Palo Alto HQ.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, PyTorch, Python, TensorFlow.
What seniority level is this role?
Pika targets lead candidates for this position.
Is this full-time or contract?
Full-time for this Research Scientist, Foundation Model (Video Generation) role at Pika.

From the original posting

About the Role

 

At Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in post-training large-scale multimodal foundation models to advance our mission of making agentic, real-time generative technology accessible and transformative for millions of creators. This is a staff and lead-level opportunity.

 

As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal post-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale.

 

What You’ll Do

 
  • Lead research and development on post-training of multimodal foundation models at scale.

  • Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.

  • Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.

  • Advance state-of-the-art techniques in diffusion, autoregressive, and other generative models for large-scale post-training and fine-tuning.

  • Identify, create, and leverage large, high-quality cross-modal datasets.

  • Bring research advancements into production-ready systems in collaboration with engineering and product teams.

  • Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.

  • Stay at the forefront of foundational model and real-time multimodal AI research.

 

What We’re Looking For

 
  • 5+ years of research experience in large-scale post-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.

  • Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).

  • Extensive hands-on experience with large-scale multimodal model design, training, and deployment.

  • Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).

  • Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.

  • Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.

  • Excellent communication and collaboration skills, and a passion for building creative enabling technology.

 

What We Offer

 
  • Competitive salary and substantial equity in a high-growth startup

  • Full health benefits + 401k matching and more

  • Collaborative, mission-driven team environment with major growth opportunities

  • Flexible on-site/remote hybrid (HQ in Palo Alto, CA)

 

About Pika

 

Pika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us and help shape the next evolution of creative technology!

 

If you are a leading researcher excited to build and scale real-time multimodal foundation models, we want to hear from you.

Similar positions

Pika
Research Scientist, Data
Pika · Palo Alto HQ
Pika
Research Intern (BS/MS/PhD)
Pika · Palo Alto HQ
Cantina
Research Scientist, Video Foundation Models
Cantina · California, United States
Genmo
Research Scientist (post-training)
Genmo · San Francisco HQ