"gpu monitoring" Jobs

215 open tech roles matching “gpu monitoring”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Kubernetes, Python, Terraform. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 215 results

Fal

Join fal as a hands-on Software Engineer to build and maintain systems for managing a large fleet of GPU servers.

Fal Remote - Global Published 6 months ago
Flexible on stack 70% coding
Fal

Join fal as a Software Engineer to build and maintain systems for managing a large fleet of GPU servers in a growing AI platform.

Fal San Francisco $180k–$250k/yr Published 6 months ago
Flexible on stack 70% coding
baseten

Join Baseten as a hands-on Operations Manager to optimize the health and utilization of our GPU fleet in a fast-growing AI company.

baseten San Francisco Published 1 week ago
Coreweave

Join CoreWeave as a Senior Software Engineer to enhance network observability for GPU cloud services in a fast-growing company.

Coreweave Sunnyvale, CA / New York City, NY / Livingston, NJ $153k–$204k/yr Published 2 weeks ago
Flexible on stack
Clockwork Systems

Join Clockwork Systems as a Senior Software Engineer to build high-performance network observability platforms for advanced distributed computing.

Clockwork Systems Onsite Palo Alto, California $140k–$210k/yr Published 4 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Graphcore

Join Graphcore as a Senior Software Engineer to optimize AI hardware and software performance in a collaborative environment.

Graphcore Gdańsk, Pomeranian Voivodeship, Poland PLN 260.4k–PLN 352.2k/yr Published 4 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Engineer to build performance insights and observability systems for AI infrastructure.

Coreweave Sunnyvale, CA / Bellevue, WA $182k–$242k/yr Published 1 month ago
Flexible on stack
Graphcore

Join Graphcore as a Senior Principal Network Engineer to design and optimize next-generation AI data center networks.

Graphcore Austin, Texas, United States Published 6 months ago
Flexible on stack
Graphcore

Lead the global teams operating Graphcore's engineering labs and data center infrastructure while ensuring reliability and efficiency.

Graphcore Austin, Texas, United States Published 1 week ago
Coram AI

Join Coram AI as a Core Engineer to build edge applications and optimize machine learning models in a fast-paced environment.

Coram AI Bangalore Published 4 months ago
Flexible on stack
Fal

Build high-performance compute environments for AI products at fal, focusing on Kubernetes and infrastructure automation.

Fal Remote - USA Published 4 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to build next-generation observability systems for large-scale AI infrastructure.

Anthropic London, UK £325k–£390k/yr Published 1 week ago
Flexible on stack
Picogrid

Join Picogrid as the first Site Reliability Engineer to enhance production reliability across cloud and edge environments.

Picogrid El Segundo, CA $170k–$195k/yr Published 1 month ago
Flexible on stack
Reflection AI

Join Reflection AI's Compute Platform team to enhance multi-cloud scheduling and GPU infrastructure in a mission-driven environment.

Reflection AI New York, NY Published 5 months ago
Triomics

Join Triomics as a Platform Engineer to build backend services and manage cloud infrastructure for processing millions of clinical documents.

Triomics New York Office Published 2 months ago
Flexible on stack AI-first team
Coreweave

Join CoreWeave as a Senior Software Engineer to build software for managing large-scale GPU data center infrastructure.

Coreweave New York, NY / Sunnyvale, CA $153k–$242k/yr Published 3 months ago
Fal

Own the reliability and security of fal's generative media model APIs in a hybrid ML Engineering/SRE role.

Fal Remote - APAC Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack