"quantization" Jobs

29 open tech roles matching “quantization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, CUDA, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 29 results

Neuralink

Join Neuralink as a Signal Processing Engineer to develop advanced DSP algorithms for brain-computer interface devices.

Neuralink Austin, Texas, United States; South San Francisco, California, United States $121k–$230.5k/yr Published 4 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focused on ML performance to optimize large language models in a fast-paced startup environment.

baseten San Francisco Published 2 years ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 3 weeks ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Twelve Labs

Lead and build a new team focused on developing Jockey Core, a reasoning LLM for video understanding at Twelve Labs.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
baseten

Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack 70% coding
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
World Labs

Join World Labs as a Performance Engineer to optimize AI models for speed and efficiency in a cutting-edge research environment.

World Labs San Francisco $200k–$300k/yr Published 4 months ago
Flexible on stack 70% coding
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 11 months ago
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Developer Advocate to shape the future of AI infrastructure through hands-on engineering and community engagement.

Fireworks AI San Francisco Published 2 weeks ago
Flexible on stack 70% coding
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack