"ml hardware accelerators" Jobs

51 open tech roles matching “ml hardware accelerators”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, CUDA. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 51 results

Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Mirendil

Join Mirendil as a staff engineer to design and optimize custom ML kernels for frontier AI research.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Anthropic

Join Anthropic as a Hardware Systems Architect to lead the design and architecture of cutting-edge AI hardware systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 1 month ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Preference Model

Join Preference Model as a Machine Learning Engineer to develop low-level reinforcement learning environments in a fast-paced startup.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Anthropic

Join Anthropic as a Software Engineer specializing in ML Networking, focusing on network infrastructure and optimization.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 11 months ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 11 months ago
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 2 days ago
Flexible on stack
Inferact

Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Anthropic

Own the end-to-end execution of custom silicon for AI systems at Anthropic, driving external partnerships and program management.

Anthropic San Francisco, CA | New York City, NY $365k–$435k/yr Published 1 month ago
World Labs

Join World Labs as a Performance Engineer to optimize AI models for speed and efficiency in a cutting-edge research environment.

World Labs San Francisco $200k–$300k/yr Published 4 months ago
Flexible on stack 70% coding
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack