"ml hardware accelerators" Jobs

158 open tech roles matching “ml hardware accelerators”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 158 results

Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working directly with hardware vendors.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Mirendil

Join Mirendil as a staff engineer to design and optimize custom ML kernels for frontier AI research.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Hardware Systems Architect to lead the design and architecture of cutting-edge AI hardware systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 1 month ago
Flexible on stack
Mythic

Join Mythic to develop the next generation of AI compilers for cutting-edge dataflow hardware.

Mythic Palo Alto, CA Published 11 months ago
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Preference Model

Join Preference Model as a Machine Learning Engineer to develop low-level reinforcement learning environments in a fast-paced startup.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Engineer to optimize GPU kernels for high-performance AI applications in a rapidly growing environment.

Coreweave Sunnyvale, CA / Bellevue, WA $182k–$242k/yr Published 1 month ago
70% coding
Anthropic

Join Anthropic as a Software Engineer specializing in ML Networking, focusing on network infrastructure and optimization.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 11 months ago
Flexible on stack
Mythic

Join Mythic to advance the MLIR ecosystem, extending high-level dialects and designing a new hardware-aware low-level dialect.

Mythic Palo Alto, CA Published 11 months ago
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 11 months ago
Coreweave
Coreweave New York, NY / Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 12 months ago
Mythic

Join Mythic as a Senior Silicon Emulation Engineer to develop and validate AI accelerators using hardware emulation platforms.

Mythic Austin, TX Published 7 months ago
AI-first team
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Kodiak Robotics

Join Kodiak Robotics as a Senior AI Infrastructure Engineer to optimize model training for autonomous technology.

Kodiak Robotics Mountain View, CA $190k–$260k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Software Engineer focused on Performance Optimization to enhance AI infrastructure efficiency and speed.

Fireworks AI San Mateo Published 1 year ago
Flexible on stack