"triton" Jobs

26 open tech roles matching “triton”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: CUDA, Triton, Python. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 26 results

Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Genmo

Join Genmo as a GPU Performance Engineer to optimize video generation models and achieve significant performance improvements.

Genmo San Francisco HQ Published 1 year ago
Flexible on stack
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 3 weeks ago
Together AI

Join Together AI as a Systems Research Engineer Intern to optimize GPU-accelerated algorithms for ML/AI applications.

Together AI San Francisco $58–$70/hr Published 5 days ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Systems Research Engineer Intern to optimize GPU programming for ML/AI applications in a collaborative environment.

Together AI San Francisco $58–$70/hr Published 5 days ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 8 months ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
baseten

Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack 70% coding
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 1 year ago
Together AI

Join Together AI as a Research Engineer to optimize large-scale training infrastructure for cutting-edge AI models.

Together AI San Francisco $200k–$290k/yr Published 1 month ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 1 week ago
Flexible on stack
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 2 months ago
Flexible on stack