"multimodal understanding" Jobs
888 open tech roles matching “multimodal understanding”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, PyTorch. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 888 results
Join Perplexity's Multimodal AI team to design and build innovative human-AI interaction systems.
Join Perplexity AI as a backend engineer to design and scale distributed systems for real-time voice interactions.
Lead the development of next-generation multimodal models at Twelve Labs, impacting thousands of customers worldwide.
Lead the architecture direction for multimodal transformers at Kodiak Robotics, focusing on AI-powered autonomous technology.
Join Cartesia as a Research Engineer to design high-quality datasets and engineer data pipelines for cutting-edge AI models.
Join Cartesia as a Researcher to enhance multimodal models through innovative post-training methods and alignment techniques.
Join Cantina as a Research Scientist to develop next-generation video foundation models and shape the future of AI-driven creativity.
Join Cantina as a Machine Learning Engineer to develop cutting-edge speech and audio generation systems in a collaborative environment.
Join Tavus as a Conversational Modelling Research Engineer to advance AI Humans through innovative multimodal conversational models.
Join Pika as a lead Research Scientist to advance real-time multimodal foundation models for creative technology.
Lead research in realtime audio understanding and human AI interaction at a pioneering AI startup.
Join Mind Robotics as a Research & Modeling Engineer to build and train core models for real-world robotic systems.
Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.
Join Mirelo AI as a Research Scientist to develop cutting-edge multimodal models for audio generation in a rapidly growing company.
Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.
Own the Generation half of Document Intelligence, transforming multimodal inputs into structured maintenance knowledge.
Own the Generation half of Document Intelligence, turning multimodal primitives into structured maintenance knowledge.
Lead a data team at Cartesia to enhance the quality of multimodal AI training data and infrastructure.
Join Descript as an Applied Research Scientist to develop multimodal understanding models for innovative AI editing features.