"gpu monitoring" Jobs

82 open tech roles matching “gpu monitoring”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Kubernetes, Python, Grafana. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 82 results

Coreweave

Join CoreWeave as a Senior Software Engineer to enhance network observability for GPU cloud services in a fast-growing company.

Coreweave Sunnyvale, CA / New York City, NY / Livingston, NJ $153k–$204k/yr Published 2 weeks ago
Flexible on stack
Clockwork Systems

Join Clockwork Systems as a Senior Software Engineer to build high-performance network observability platforms for advanced distributed computing.

Clockwork Systems Onsite Palo Alto, California $140k–$210k/yr Published 4 months ago
Flexible on stack
Graphcore

Join Graphcore as a Senior Software Engineer to optimize AI hardware and software performance in a collaborative environment.

Graphcore Gdańsk, Pomeranian Voivodeship, Poland PLN 260.4k–PLN 352.2k/yr Published 4 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Engineer to build performance insights and observability systems for AI infrastructure.

Coreweave Sunnyvale, CA / Bellevue, WA $182k–$242k/yr Published 1 month ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Software Engineer to build software for managing large-scale GPU data center infrastructure.

Coreweave New York, NY / Sunnyvale, CA $153k–$242k/yr Published 3 months ago
Fal

Own the reliability and security of fal's generative media model APIs in a hybrid ML Engineering/SRE role.

Fal Remote - APAC Published 2 months ago
Flexible on stack
Fundamental

Join Fundamental as a DevOps Engineer to tackle technical challenges in AI infrastructure for enterprise decision-making.

Fundamental Europe Published 9 months ago
krea.ai

Join Krea as an ML Researcher to train diffusion models for image and video generation in a creative AI-focused environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Together AI

Join Together AI as a Senior Software Engineer to design and implement a scalable observability platform for our generative AI lifecycle.

Together AI San Francisco $200k–$280k/yr Published 10 months ago
Flexible on stack
Ambient

Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced AI models.

Ambient Redwood City Published 2 months ago
Flexible on stack 70% coding
Twelve Labs

Lead the Construction team at Twelve Labs to enhance engineering execution and drive complex platform projects.

Twelve Labs Seoul, South Korea Published 1 day ago
Flexible on stack Heavy meetings
Datadog

Lead the applied science direction for GenSim, building realistic simulated environments for Datadog's AI agents.

Datadog Paris, France Published 1 week ago
Flexible on stack
Fundamental

Join Fundamental as an MLOps Engineer to tackle technical challenges in AI and transform enterprise decision-making.

Fundamental Europe Published 1 month ago
Flexible on stack
TRM Labs

Join TRM Labs as a Senior Frontend Platform Engineer to build high-performance visualization systems for blockchain investigations.

TRM Labs North America $210k–$230k/yr Published 1 month ago
Flexible on stack
Armada

Lead the global AI infrastructure supply chain as an IT Hardware Procurement Manager at Armada, focusing on high-performance computing.

Armada India (Remote) Published 3 months ago
Nomic AI

Join Nomic AI as a Senior Platform Engineer to own and scale our infrastructure stack in the AEC industry.

Nomic AI New York HQ Published 2 months ago
Flexible on stack
TRM Labs

Join TRM Labs as a Senior Frontend Platform Engineer to build high-performance visualization systems for blockchain investigations.

TRM Labs South America Published 6 months ago
Flexible on stack
Databricks

Join Databricks as a Senior Software Engineer to build and scale a managed GPU training platform for AI models.

Databricks Mountain View, California; San Francisco, California $160k–$225k/yr Published 3 months ago
Flexible on stack 70% coding
SpaceX

Join SpaceX as a Sr. HPC Systems Engineer to manage HPC clusters and support engineering teams in a fast-paced environment.

SpaceX Hawthorne, CA $165k–$230k/yr Published 2 months ago
Flexible on stack