Join Perplexity as a Search Crawler Analyst to enhance our crawling and storage system through data analysis and engineering.
Posted by employer 11 hours ago
First seen on Joblaze 4 hours ago
Last verified on the company career page 4 hours ago
Skills & Technologies
What you'll build
Must have
Nice to have
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the individual will focus on diagnosing and resolving quality issues within Perplexity's complex crawling pipeline while also enhancing the overall search and answer systems. Key skills include strong coding abilities at a mid-level backend engineer standard, experience in data analysis, and a background in machine learning model training. This position is well-suited for candidates with over four years of relevant experience, particularly those familiar with web crawling or indexing processes. The team operates at the intersection of data analysis and engineering, emphasizing continuous improvement.
Joblaze insights
Quick facts
From the original posting
The internet is vast, containing trillions of URLs. Perplexity’s crawling and storage system is complex and has multiple stages (URL discovery, crawling, parsing, indexing). Each stage offers many opportunities for improvement and room for intricate bugs. You’ll work at the intersection of data analysis and engineering — designing metrics, building data pipelines, and improving the quality of our search and answer systems.
Responsibilities
Find and diagnose quality issues in our crawling pipeline
Train small models that optimize particular aspects of the pipeline (e.g. parsing quality)
Build datasets for model training, including LLM-as-a-judge labeling pipelines
Improve page selection algorithms for indexing
Design and analyze experiments to validate improvements
Qualifications
4+ years of experience as a data analyst, ML engineer, or in a related role
Strong coding skills: you should be able to write production-grade code at the level of a mid-level backend engineer
Experience designing metrics from scratch
Experience training ML models that shipped to production with measurable metric improvements
Nice to have
Direct experience working on web crawling or indexing pipelines