"gpu monitoring" Jobs

29 open tech roles matching “gpu monitoring”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Kubernetes, Python. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 29 results

Datadog

Contribute to GPU Monitoring capabilities within the Datadog Agent while working on eBPF infrastructure.

Datadog Portugal, Remote Published 2 months ago
Flexible on stack
Datadog

Contribute to GPU Monitoring capabilities within the Datadog Agent while working with eBPF and Linux kernel infrastructure.

Datadog Denmark, Remote; France, Remote; Germany, Remote; Ireland, Remote; Italy, Remote; Poland, Remote; Spain, Remote; Sweden, Remote; Switzerland, Remote; United Kingdom, Remote Published 2 months ago
Flexible on stack
Datadog

Own the go-to-market strategy for Infrastructure Monitoring for AI and containerized workloads at Datadog.

Datadog New York, New York, USA $123k–$164k/yr Published 5 days ago
Together AI

Join Together AI as a Senior Software Engineer to design and implement a scalable observability platform for our generative AI lifecycle.

Together AI San Francisco $200k–$280k/yr Published 8 months ago
Flexible on stack
Reddit

Join Reddit as a Staff Machine Learning Engineer to enhance ML efficiency and drive performance improvements across the platform.

Reddit Remote - The Netherlands Published 1 month ago
Flexible on stack
Reddit

Join Reddit as a Staff Machine Learning Engineer to enhance ML efficiency and drive performance improvements across the company's ML ecosystem.

Reddit Remote - United Kingdom Published 1 month ago
Flexible on stack
Datadog

Lead engineering for Cloud Observability at Datadog, managing a team of ~40 engineers in a hybrid work environment.

Datadog Boston, Massachusetts, USA; New York, New York, USA $296k–$370k/yr Published 1 week ago
Together AI

Join Together AI as a Manager of Infrastructure Strategy & Operations to drive analytical decisions in scaling compute infrastructure.

Together AI San Francisco $220k–$260k/yr Published 1 month ago
Together AI

Join Together AI as an AI Infrastructure Engineer (SRE) to ensure the reliability and scalability of user-facing services.

Together AI Bangalore, India Published 2 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Staff+ Software Engineer to build production systems for capacity engineering in a hybrid work environment.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 2 weeks ago
Flexible on stack 70% coding
Datadog

Join Datadog as a Senior Applied Scientist to design and optimize AI models for real-time anomaly detection in security products.

Datadog Paris, France Published 3 weeks ago