Join Exa's Infrastructure Team to build massive-scale ML systems that redefine how AI consumes information.
Posted by employer 1 year ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
Skills & Technologies
AI in the day-to-day
AI can 10x our velocity but also 10x the strain on our systems, and that challenge excites you.
Requirements
Not disclosed in this posting: compensation, years of experience.
Benefits
Joblaze summary
In this role, the Software Engineer for Infrastructure at Exa focuses on developing and scaling the foundational systems that support the company's ambitious AI-driven search engine. Key responsibilities include orchestrating GPU-based training and inference across multiple cloud environments, as well as creating advanced CI and build tools. This position is ideal for engineers with a strong background in distributed systems and a passion for automation, particularly those who thrive in fast-paced, innovative settings. Exa's infrastructure team plays a crucial role in enabling rapid development and deployment of cutting-edge AI technologies.
Joblaze insights
Quick facts
From the original posting
Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.
Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
Our Infrastructure Team builds the underlying tooling and infrastructure that powers all Exa's systems. Basically, infra engineers build the machine that builds the machine so that we can move as fast as possible as an engineering org. That could mean building GPU cluster orchestration in Kubernetes, map-reduce batchjobs on Ray, or the best observability tooling in the world.
You’re obsessed with scale and complex distributed systems
You have an aversion to ClickOps and would always build an automation
You max out OSS models’ GPU flops utilization just for the challenge
You play with Arch or NixOS as a personal driver for fun or ideology
You’re passionate and discerning about AI tooling — AI can 10x our velocity but also 10x the strain on our systems (CI, etc) and that challenge excites you
Scale infrastructure to process the whole web on GPUs cost efficiently
Orchestrate multi-region, multi-cloud inference and training on GPUs
Ship a new version of our LLM gateway
Make the world's most advanced build cache with Nix
Build custom CI and code infrastructure that scales for our agent fleet
Automate software maintenance and improvements for the whole company
Location: This is an in-person opportunity in San Francisco.
Visas: We're happy to sponsor international candidates (e.g., STEM OPT, OPT, H1B, O1, E3). While we cannot guarantee your visa, we have historically been successful in sponsoring candidates from all over the world. If you receive an offer, our team will work hard to get you a visa.
Benefits: We offer premium healthcare benefits (medical, dental, vision), fertility benefits, 16 weeks of fully paid parental leave for all new parents, and a monthly wellness stipend to all of our employees.
Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.