← Back to results

Research Engineer, Content Understanding

Join Exa as a Research Engineer to advance AI-driven content understanding for a revolutionary search engine.

Location
San Francisco, California, United States
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 3 days ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Exa → Save job Scanned from exa.ai

What you'll build

  • Make parsing work on web pages
  • Teach a model to judge page quality
  • Work on credibility and misinformation
  • Decide whether two documents are semantically the same
  • Build classification and extraction at web scale

Must have

  • Graduate-level ML experience
  • Build a transformer from scratch in PyTorch
  • Experience with large-scale datasets
  • Comfortable with undefined ground truth problems
  • Care about finding high quality knowledge

AI in the day-to-day

We build massive-scale ML systems that will define the way the new AI world consumes information.

Requirements

Experience
2+ years
Education
Master's degree

Not disclosed in this posting: compensation, work arrangement, visa sponsorship.

Joblaze summary

In this role, the research engineer focuses on enhancing the search architecture by developing methods to accurately parse and classify web pages, ensuring high-quality search results. Proficiency in machine learning, particularly in building transformers with PyTorch and handling large-scale datasets, is essential. This position is suited for individuals with advanced degrees in ML or strong undergraduates who are comfortable tackling ambiguous problems related to content credibility and misinformation. Exa's innovative environment encourages flexibility in project selection based on individual skills and interests.

Joblaze insights

  • Listed yesterday — first seen on Joblaze October 3, 2026. Last confirmed on Exa's careers page October 3, 2026.
  • Machine Learning appears in 11.1% of 460 comparable mid ai/ml roles in United States; document understanding appears in 0.2% of 460 comparable mid ai/ml roles in United States.

Quick facts

How much experience is required?
At least 2 years of relevant experience for this Research Engineer, Content Understanding role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Machine Learning, PyTorch, data processing, document understanding.
What seniority level is this role?
Exa targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Research Engineer, Content Understanding role at Exa.

From the original posting

Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.

 

Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.

 

As a backend engineer, you'd play a critical role in our search architecture. We're pretty flexible on what projects people work on based on their skills and interests.

Search quality is bounded by what we understand about a page. Before anything can be retrieved, something has to work out what the page actually says. That means parsing it into the parts that are content and the parts that are furniture, classifying what kind of page it is and what it is about, telling whether the page is usable at all, extracting when it was published, judging how good it is and whether it can be trusted, and working out whether it says anything that a page we already have does not. All of this has to work on every page on the web, in every language, in every shape the web comes in.

Some of this is classic document understanding. Some of it is much more open. Credibility and misinformation, AI-generated and machine-spun content, and pages written to be found rather than read are all unsolved, and search results are only as trustworthy as our answers to them.

We are looking for a research engineer to work on this. There is a lot of room to do it well.

Desired Experience

  • Graduate-level ML experience (Master’s or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad

  • You can build a transformer from scratch in PyTorch, and you have trained models that then had to be cheap enough to run everywhere

  • You like building large-scale datasets and living in the data. Most of the wins here are in the supervision rather than the architecture

  • You are comfortable with problems where the ground truth does not exist yet and defining it is part of the job

  • You care about the problem of finding high quality knowledge and recognize how important this is for the world

Example Projects

  • Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it

  • Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it

  • Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler

  • Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately

  • Build classification and extraction that is accurate at web scale and cheap enough to run on all of it

  • Design the supervision for something nobody has labels for, and find out whether it is learnable at all

  • Trace a bad search result back to the page-level prediction that caused it, and fix it at the source

Standard company text repeated across Exa's postings is omitted here.

Similar positions

Exa
Research, ML
Exa · San Francisco, California
Exa
Research Engineer, Index Intelligence
Exa · San Francisco, California, United States
Exa
Research Engineer, Generalist
Exa · San Francisco, California
Exa
Research, ML
Exa · Singapore
Exa
Software Engineer, Backend
Exa · San Francisco, California