Posted by employer 6 months ago
First seen on Joblaze 5 months ago
Last verified on the company career page 3 hours ago
Not disclosed in this posting: compensation, visa sponsorship.
Quick facts
- Is the Lead/Manager AI Infra Systems Engineering Team (Amsterdam) role remote?
- No — this is an on-site role in Amsterdam .
- How much experience is required?
- At least 7 years of relevant experience for this Lead/Manager AI Infra Systems Engineering Team (Amsterdam) role.
- Where is the role based?
- Together AI is hiring for this position in Amsterdam .
- What's the tech stack?
- Joblaze extracted these technologies from the posting: Ansible, Cloud Services, Kubernetes, Terraform.
- What seniority level is this role?
- Together AI targets lead candidates for this position.
- Is this full-time or contract?
- Full-time for this Lead/Manager AI Infra Systems Engineering Team (Amsterdam) role at Together AI.
From the original posting
About the Role
Lead a team of AI Infra (Systems Engineers) at Together based out of our office in Amsterdam, you and the SRE team are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebase.
You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems.
Responsibilities
- Be on an on-call (PagerDuty) rotation to respond to incidents that impact availability
- Manage, develop and coach the SRE Team.
- Build and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent users
- Build monitoring systems to ensure the highest quality service for our customers
- Design and implement operational processes (such as deployments and upgrades)
- Debug production issues across all services and levels of the stack
- Identify improvements for the product architecture from the reliability, performance and availability perspectives
- Plan the growth of Together AI’s infrastructure
Requirements
- 7+ years of professional SRE or related experience
- Ideally 2 years as a Lead SRE
- Bachelor's degree in Computer Science or a related field or equivalent work experience
- Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes
- Proficiency in programming/scripting languages
- Direct experience in monitoring and observability practices
- Advanced knowledge of cloud services
- Ability to thrive in a collaborative environment involving different stakeholders and subject matter experts
Please see our privacy policy at https://www.together.ai/privacy
Standard company text repeated across Together AI's postings is omitted here.