"gpu monitoring" Jobs

65 open tech roles matching “gpu monitoring”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Kubernetes, Python, Terraform. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 65 results

TRM Labs

Join TRM Labs as a Senior Frontend Platform Engineer to build high-performance visualization systems for blockchain investigations.

TRM Labs North America $210k–$230k/yr Published 1 month ago
Flexible on stack
Mithril

Join Mithril as a Site Reliability Engineer to enhance the stability and performance of our global GPU orchestration platform.

Mithril Palo Alto / San Francisco Bay Area $170k–$230k/yr Published 4 months ago
Flexible on stack 70% coding
TRM Labs

Join TRM Labs as a Senior Frontend Platform Engineer to build high-performance visualization systems for blockchain investigations.

TRM Labs South America Published 6 months ago
Flexible on stack
Mithril

Join Mithril as a Software Engineer to build and maintain backend systems for a cloud-native AI infrastructure platform.

Mithril Palo Alto / San Francisco Bay Area $170k–$230k/yr Published 4 months ago
Flexible on stack
Databricks

Join Databricks as a Senior Software Engineer to build and scale a managed GPU training platform for AI models.

Databricks Mountain View, California; San Francisco, California $160k–$225k/yr Published 3 months ago
Flexible on stack 70% coding
Omnifold

Lead the infrastructure team at Omnifold, focusing on AI model training and deployment in a fast-paced startup environment.

Omnifold San Francisco HQ Published 6 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to drive the architecture of a managed GPU training platform for large-scale AI models.

Databricks Mountain View, California; San Francisco, California $190k–$265k/yr Published 3 months ago
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Together AI

Design and maintain high-performance networks as a Senior Network Engineer at Together AI, contributing to cutting-edge AI infrastructure.

Together AI San Francisco $190k–$270k/yr Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Manager of Infrastructure Strategy & Operations to drive analytical decisions in scaling compute infrastructure.

Together AI San Francisco $220k–$260k/yr Published 3 months ago
Exa

Join Exa as a Security Engineer to build protective systems for massive-scale AI infrastructure in San Francisco.

Exa San Francisco, California Published 2 weeks ago
AI-first team
Databricks

Develop and run the research stack that powers Databricks AI Research, enabling rapid large-scale experiments.

Databricks New York City, New York; San Francisco, California $199k–$270k/yr Published 4 months ago
Flexible on stack
Databricks
Databricks New York City, New York; San Francisco, California $190k–$270k/yr Published 4 months ago
Mithril

Join Mithril as a Software Engineer to build and maintain backend systems for a cutting-edge AI infrastructure platform.

Mithril Palo Alto / San Francisco Bay Area $170k–$230k/yr Published 4 months ago
Flexible on stack
Inferact

Join Inferact as a Site Reliability Engineer to enhance the reliability and performance of AI inference systems at scale.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack AI-first team
Fal

Join fal as a Senior Site Reliability Engineer to enhance the reliability of our generative media infrastructure in San Francisco.

Fal San Francisco Published 6 months ago
baseten

Lead a team of cloud platform engineers to build scalable and reliable infrastructure for AI products at Baseten.

baseten San Francisco Published 3 months ago
Flexible on stack Heavy meetings
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack