"cuda" Jobs
130 open tech roles matching “cuda”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: CUDA, Python, Kubernetes. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 130 results
Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.
Join Genmo as a GPU Performance Engineer to optimize video generation models and achieve significant performance improvements.
Join CoreWeave as a Senior Engineer to optimize GPU kernels for high-performance AI applications in a rapidly growing environment.
Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.
Join Fireworks AI as a Software Engineer focused on Performance Optimization to enhance AI infrastructure efficiency and speed.
Join Sarvam AI as a Senior Performance Engineer to optimize GPU kernels for high-performance ML systems.
Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.
Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.
Join Fundamental as a Senior Applied Research Engineer to tackle technical challenges in AI model development for enterprise decision-making.
Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working directly with hardware vendors.
Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.
Join Fireworks AI as a Member of Technical Staff to design and build systems infrastructure for AI workloads at scale.
Develop real-time DSP algorithms for space-based communications at Cowboy Space Corporation, a pioneering energy startup.
Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.
Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.
Related searches