"low precision inference" Jobs

51 open tech roles matching “low precision inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, PyTorch, CUDA. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 51 results

Applied Intuition

Join Applied Intuition as an AI Performance Engineer to optimize large-scale machine learning workloads in a collaborative environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.

Perplexity AI London Published 5 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working directly with hardware vendors.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Staff Software Engineer to lead the development of a Kubernetes-native inference platform for AI workloads.

Coreweave Sunnyvale, CA / Bellevue, WA $188k–$275k/yr Published 4 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
World Labs

Join World Labs as a Performance Engineer to optimize AI models for speed and efficiency in a cutting-edge research environment.

World Labs San Francisco $200k–$300k/yr Published 4 months ago
Flexible on stack 70% coding
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 1 year ago
Fireworks AI

Join Fireworks AI as a Software Engineer focused on Performance Optimization to enhance AI infrastructure efficiency and speed.

Fireworks AI San Mateo Published 1 year ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 2 years ago
krea.ai

Join Krea as an ML Researcher to train diffusion models for image and video generation in a creative AI-focused environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
baseten

Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack 70% coding
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago