Join Together AI as a Staff Software Engineer to build systems that automate infrastructure management for AI clusters.
Posted by employer 1 month ago
First seen on Joblaze 1 month ago
Last verified on the company career page 14 hours ago
Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.
Joblaze summary
In this role, the Staff Software Engineer focuses on developing systems that automate the lifecycle management of hardware for AI inference clusters, enabling teams to provision resources with a single API call. Proficiency in languages like Go, Python, or Rust, along with experience in durable workflow orchestration and event-driven systems, is essential. This position is ideal for seasoned engineers with a product mindset who have previously built internal platforms or APIs. Together AI, a research-driven company, emphasizes innovation in AI infrastructure and values a collaborative team environment.
Joblaze insights
Quick facts
From the original posting
We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle — turning racks of GPUs into running inference clusters without a human touching a runbook. The Research and Inference team is your customer: today they file tickets and wait; the target state is that they issue a single API call to stand up, scale, or tear down a cluster, and the system takes care of the rest. The platform is manifest-driven such that teams declare the desired state of a cluster or host — shape, topology, software stack — and the system is responsible for reconciling reality to that manifest, continuously, through every stage of its lifecycle. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry a piece of hardware or a cluster from one state to the next—taking it from bare metal to a fully functioning AI cluster for training or inference.
You'll write production code which is typed, tested, versioned, and deployed through CI/CD that models infrastructure state and reconciles it, the same way a Kubernetes controller reconciles a cluster's desired state. Success looks like eliminating manual provisioning work, not documenting it better.
A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production.
Core requirements (all levels):
Nice to have:
Please see our privacy policy at https://www.together.ai/privacy.
Standard company text repeated across Together AI's postings is omitted here.