"multimodal understanding" Jobs
392 open tech roles matching “multimodal understanding”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, PyTorch. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 392 results
Join Cantina as a Research Scientist to develop next-generation video foundation models and shape the future of AI-driven creativity.
Join Cantina as a Machine Learning Engineer to develop cutting-edge speech and audio generation systems in a collaborative environment.
Lead research in realtime audio understanding and human AI interaction at a pioneering AI startup.
Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.
Join Mirelo AI as a Research Scientist to develop cutting-edge multimodal models for audio generation in a rapidly growing company.
Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.
Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.
Join Reddit as a Senior ML Engineer to develop advanced embedding models for advertising in a fully remote role.
Join Reddit as a Senior ML Engineer to develop advanced embedding models for advertising in a fully remote role.
Join Ambient.ai as a Senior Applied Research Scientist to develop cutting-edge foundation models for computer vision in a hybrid work environment.
Join Sarvam AI as a researcher to develop vision-language models that impact AI applications in India.
Join Cartesia as a Researcher in London to advance AI through innovative neural network architecture design.
Join Twelve Labs as a Senior AI Engineer to build the integration layer for their multimodal video AI system.
Own the interfaces for Model APIs and Developer Experience at Together AI, focusing on multimodal APIs and ecosystem compatibility.
Join Protege as a Machine Learning Researcher to lead the evaluation and optimization of audio data quality for AI training.
Lead research in audio-visual avatar generation at Tavus, shaping the future of human-AI interaction.