"interpretability" Jobs

12029 open tech roles matching “interpretability”, taken straight from company career pages — not reposted from other job boards. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 12029 results

Anthropic

Join Anthropic as a Staff Software Engineer to drive reinforcement learning efforts and design systems for coding capabilities.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$625k/yr Published 4 days ago
Anthropic
Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY $320k–$485k/yr Published 3 months ago
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 2 months ago
Flexible on stack
Anthropic
Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Washington, DC; San Francisco, CA | New York City, NY $230k–$270k/yr Published 4 months ago
Anthropic

Join Anthropic as a Safeguards Analyst to build enforcement workflows for AI systems against misuse and ensure user safety.

Anthropic Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC $285k–$330k/yr Published 2 weeks ago
Anthropic

Lead the Model Exploitation & Fraud team at Anthropic to combat large-scale exploitation of AI systems.

Anthropic San Francisco, CA $375k–$455k/yr Published 2 weeks ago
Anthropic

Build evaluation infrastructure for AI safety systems at Anthropic, focusing on real-world misuse detection.

Anthropic San Francisco, CA | New York City, NY $320k–$485k/yr Published 1 month ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY $300k–$405k/yr Published 6 months ago
Anthropic

Join Anthropic as a Data Engineer to build data infrastructure that ensures AI systems are safe and beneficial.

Anthropic San Francisco, CA | New York City, NY $1–$2/yr Published 1 week ago
Flexible on stack
Inflection AI

Lead the architecture and technical direction of agentic AI systems at Inflection AI, building production AI agents.

Inflection AI Palo Alto, California, United States $400k–$550k/yr Published 4 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Research Scientist focusing on measuring and understanding recursive self-improvement in AI systems.

Anthropic San Francisco, CA | New York City, NY $350k–$850k/yr Published 4 days ago