Own the infrastructure and CI/CD pipelines for a cutting-edge AI lab focused on chip design.
Posted by employer 1 day ago
First seen on Joblaze 2 hours ago
Last verified on the company career page 2 hours ago
Skills & Technologies
What you'll build
Must have
Nice to have
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the DevOps Engineer at Ricursive Intelligence is responsible for managing and enhancing the infrastructure that supports advanced AI research, including cloud platforms and CI/CD pipelines. Key skills include expertise in infrastructure-as-code, continuous integration, and observability tooling, with a focus on security and scalability. This position is ideal for someone with significant experience in infrastructure engineering, particularly in fast-paced environments, who can adapt to evolving technical demands. The team is composed of top-tier talent from leading tech companies, contributing to a high-performance culture.
Joblaze insights
Quick facts
From the original posting
ABOUT THE ROLE
Ricursive's research runs on infrastructure that has to keep pace with the research itself: ML training and evaluation workloads, EDA tool flows, and a fast-growing team that needs everything from cloud environments to developer workflows to just work. This role owns that foundation — the pipelines, platforms, and systems that let a small team move like a much larger one.
You will own our infrastructure-as-code (IaC), continuous integration & deployment (CI/CD), and cloud platform end-to-end that keeps the lab running day to day. Additionally, you will be collaborating closely with the team to build out the right observability stack for their needs while working around environment security limitations.
WHAT YOU WILL DO
• Design, build out and extend our self-managed cloud platform with Terraform, setting the IaC patterns the rest of the team builds on.
• Own our platform deployments spanning various environments day to day, including performance, security, reliability, scalability, and adapting to evolving business requirements.
• Design, implement, and operate CI/CD pipelines in GitHub Actions for mission-critical repositories, working within strict security and deployment restrictions.
• Architect scalable AI tooling and developer experience workflows across multiple distinct user environments, working closely with our physical design (PD) engineers.
• Build out the observability stack with the team, covering pipeline health, application-level metrics, and ML workloads.
MINIMUM QUALIFICATIONS
• BS in CS, CE, EE, or a closely related technical field, or equivalent practical experience.
• 4+ years of hands-on infrastructure/platform engineering, with ownership of systems others depend on.
• Owned production IaC architecture including maintenance and new features, not just consumed modules.
• Designed and implemented production CI/CD pipelines that build and ship reproducible artifacts, with attention to performance, scalability, and security.
• Prior experience with observability tooling for monitoring pipeline health as well as application-level metrics.
PREFERRED QUALIFICATIONS
• Hands-on experience with GCP, Kubernetes, and GitHub Actions, including custom runner setups.
• Experience running ML infrastructure for training and evaluation workloads, including GPU/TPU compute.
• Familiarity with LLM observability tooling: tracing, evaluations, and cost & latency monitoring.
• Security depth: dependency supply-chain hardening, OIDC-based auth, least-privilege secrets, and compliance work such as SOC 2 or penetration testing.
• Early-stage startup experience: built infrastructure from zero or near-zero.
Standard company text repeated across Ricursive Intelligence's postings is omitted here.