Lead complex infrastructure projects to support AI research at Anthropic with a focus on reliability and scalability.
Posted by employer 6 hours ago
First seen on Joblaze 2 hours ago
Last verified on the company career page 2 hours ago
Joblaze summary
In this role, the Staff+ Software Engineer will collaborate closely with research teams to develop and maintain the infrastructure essential for training and deploying AI models. Key skills include expertise in large-scale distributed systems, proficiency in programming languages like Python or Go, and experience with cloud technologies such as Kubernetes and AWS. This position is ideal for seasoned engineers with a strong background in technical leadership and project management. Anthropic emphasizes a collaborative environment, focusing on impactful AI research.
Quick facts
- Is the Staff+ Software Engineer, Research Systems Engineering role remote?
- It's hybrid — Anthropic expects some on-site time in San Francisco, CA, United States.
- What's the salary range?
- Anthropic lists $320,000–$485,000 for this role.
- How much experience is required?
- At least 10 years of relevant experience for this Staff+ Software Engineer, Research Systems Engineering role.
- Where is the role based?
- Anthropic is hiring for this position in San Francisco, CA, United States.
- What's the tech stack?
- Joblaze extracted these technologies from the posting: AWS, GCP, Go, Java, Kubernetes, Python.
- Does Anthropic sponsor work visas for this role?
- Yes — the posting indicates visa sponsorship is available for the right candidate.
- What seniority level is this role?
- Anthropic targets staff-level candidates for this position.
- Is this full-time or contract?
- Full-time for this Staff+ Software Engineer, Research Systems Engineering role at Anthropic.
From the original posting
About Anthropic
About the role
Anthropic's Infrastructure organization builds and operates the distributed systems that train, serve, and secure our AI models. Every other team at Anthropic depends on these systems.
In this role, you'll work directly with our Research teams to build the reliable, scalable, performant infrastructure our research depends on. This is a cross-functional systems team: you'll scope and lead complex, multi-month infrastructure projects and resolve the performance and scalability bottlenecks that limit how fast we can grow.
Key responsibilities
- Independently scope and lead complex, multi-month infrastructure projects, from an ambiguous starting point through to a production system
- Build deep partnerships with researchers and Research teams to understand their needs and deliver for them
- Mentor other engineers and help raise the technical bar for the team
- Build alignment on technical direction across multiple teams, working through ambiguous problem spaces
- Take ownership of the reliability, scalability, and security of the systems you build as usage and complexity grow
- Lead the improvement of operational processes across Infrastructure, such as incident response, postmortems, and on-call rotations, so the team learns from every incident
Minimum qualifications
- Experience designing, building, and operating large-scale distributed systems or infrastructure in production
- Experience independently scoping and delivering complex, ambiguous, multi-month technical projects
- Prior experience as a technical lead or mentor for other engineers
- Experience making architectural decisions that other engineers and teams build on top of
- Strong software engineering fundamentals and proficiency in at least one programming language (e.g., Python, Rust, Go, Java)
- Experience with modern cloud infrastructure (e.g., Kubernetes, infrastructure as code, AWS, GCP)
- Strong written and verbal communication skills, with experience building alignment across multiple teams or stakeholders
Preferred qualifications
- 10+ years of software engineering experience, not including internships
- Experience with machine learning infrastructure (e.g., GPUs, TPUs, Trainium) and associated networking infrastructure (e.g., NCCL)
- Low-level systems experience (e.g., Linux kernel tuning, eBPF)
- Experience applying security or privacy engineering best practices
The annual compensation range for this role is listed below.
Annual Salary:
$320,000—$485,000 USD
Logistics
Standard company text repeated across Anthropic's postings is omitted here.