"cluster management" Jobs
833 open tech roles matching “cluster management”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Kubernetes, Python, Go. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 833 results
Join SpaceX as a Sourcing Manager to lead the strategy for rack-level cooling and mechanical components in a high-velocity environment.
Join Cresta as a Senior Infrastructure Engineer/SRE to design and build core infrastructure for AI-driven customer experiences.
Lead full-stack delivery of data center projects across Europe, ensuring rapid deployment of NVIDIA clusters.
Join Graphcore as a Staff Engineer to develop critical system management interfaces for AI solutions in a collaborative environment.
Join Okta as a Staff SRE to architect and manage Kubernetes platforms on AWS, ensuring high availability and performance.
Design and automate infrastructure for robotics systems at Dexterity, focusing on scalability and security.
Join PlanetScale as a Software Engineer to design and build the control plane for Neki, a sharded PostgreSQL product.
Join Clockwork Systems as a Senior Software Engineer to build high-performance network observability platforms for advanced distributed computing.
Drive the strategy and evolution of the Calico platform as a Principal Product Manager in a hybrid role based in the San Francisco Bay Area.
Lead sourcing strategy and supplier management for server mechanicals and cooling at SpaceX, driving innovation and cost efficiency.
Lead a team of engineers to enhance MongoDB Atlas features while collaborating with cross-functional partners in a hybrid role based in Dublin.
Lead the Compute Platform team at Reflection AI, focusing on multi-cloud scheduling and GPU deployments while mentoring a team of systems engineers.
Build a real-time AI copilot for live commerce hosts to enhance their performance during shows.
Lead the design and management of next-gen data centers at SpaceX, driving innovation in high-stakes environments.
Join Cloudflare as a Network Reliability Engineer to enhance network resilience and automate operational tasks in a hybrid work environment.
Lead a team of program managers to oversee technical programs and ensure successful delivery of AI infrastructure projects.
Manage international cloud sourcing and supplier relationships to optimize GPU compute capacity and costs at Together AI.