Join Exa as a Web Crawler engineer to build a massive-scale web crawling system in San Francisco.
Posted by employer 1 year ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
Skills & Technologies
AI in the day-to-day
We build massive-scale infra to crawl the entire web and train state-of-the-art embedding models.
Requirements
Not disclosed in this posting: compensation, years of experience.
Benefits
Joblaze summary
In the role of Software Engineer focused on web crawling at Exa, the individual will develop and optimize large-scale web crawlers capable of processing millions of pages daily. Proficiency in high-performance programming languages like C++ or Rust, along with familiarity in TypeScript and modern web technologies, is essential for success. This position is ideal for experienced engineers who are passionate about enhancing information retrieval systems and are eager to tackle complex challenges in a cutting-edge AI environment.
Joblaze insights
Quick facts
From the original posting
Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.
Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
As a Web Crawler engineer, you'd be responsible for crawling the entire web. Basically build Google-scale crawling!
You have extensive experience building and scaling web crawlers, or would be excited to ramp up very quickly
You have experience with some high performance language (C++, Rust, etc.)
You are familiar with TypeScript, Playwright, modern web design, CDP (Chrome DevTools Protocol)
You’re comfortable optimizing a system to an exceptional degree
You care about the problem of finding high quality knowledge and recognize how important this is for the world
Build a distributed crawler that can handle 100M+ pages per day
Optimize crawl politeness and rate limiting across thousands of domains
Design systems to detect and handle dynamic content, JavaScript rendering, and anti-bot measures
Create intelligent crawl scheduling and prioritization algorithms for maximum coverage efficiency
Location: This is an in-person opportunity in San Francisco.
Visas: We're happy to sponsor international candidates (e.g., STEM OPT, OPT, H1B, O1, E3). While we cannot guarantee your visa, we have historically been successful in sponsoring candidates from all over the world. If you receive an offer, our team will work hard to get you a visa.
Benefits: We offer premium healthcare benefits (medical, dental, vision), fertility benefits, 16 weeks of fully paid parental leave for all new parents, and a monthly wellness stipend to all of our employees.
Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.