Join Exa as a Data Engineer to build massive-scale data infrastructure for a revolutionary AI-powered search engine.
Posted by employer 8 months ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
Skills & Technologies
AI in the day-to-day
Engineers will build massive-scale ML systems that define how AI consumes information.
Requirements
Not disclosed in this posting: compensation, years of experience.
Benefits
Joblaze summary
In this role, the Software Engineer will focus on designing and building the data infrastructure that supports Exa's ambitious AI-driven search engine, handling vast amounts of web data and real-time indexing. Key skills include expertise in lakehouse architectures, experience with large-scale distributed data processing, and familiarity with streaming data systems. This position is ideal for seasoned engineers with a strong background in data systems who thrive in high-autonomy environments. Exa's innovative approach and significant funding position it as a leader in the evolving AI landscape.
Joblaze insights
Quick facts
From the original posting
Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.
Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
As a Data Engineer, you'll architect and build the data infrastructure that powers everything we do—from crawling billions of pages to training our embedding models to serving real-time search. You'll have enormous autonomy in designing systems that scale to hundreds of petabytes. If you've ever wanted to build data pipelines at a scale that most companies only dream about, this is your chance.
Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them
Experience building and operating large-scale distributed data processing pipelines
Hands-on experience with streaming data systems (Kafka, Flink, or similar)
Familiarity with Ray, Spark, or ClickHouse at production scale
An obsessive focus on reliability and building systems that don't page you at 3am
Experience with Lance or other vector-native storage formats
Background in GPU-accelerated data processing (RAPIDS, cuDF)
Design a lakehouse architecture that handles 100+ PB of web crawl data
Build streaming pipelines that process billions of documents per day for real-time indexing
Architect the data layer for our embedding training infrastructure on Ray
Scale our ClickHouse deployment to handle analytical queries across petabytes of search logs
Location: This is an in-person opportunity in San Francisco.
Visas: We're happy to sponsor international candidates (e.g., STEM OPT, OPT, H1B, O1, E3). While we cannot guarantee your visa, we have historically been successful in sponsoring candidates from all over the world. If you receive an offer, our team will work hard to get you a visa.
Benefits: We offer premium healthcare benefits (medical, dental, vision), fertility benefits, 16 weeks of fully paid parental leave for all new parents, and a monthly wellness stipend to all of our employees.
Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.