← Back to results

AI Evaluation Infrastructure Engineer

Build infrastructure for high-quality AI evaluation at Block, enabling faster and more reliable AI product development.

Location
Bay Area, CA, United States
Compensation
$263.6k–$395.4k/yr
Level
mid
Type
full time · Hybrid

Posted by employer 3 days ago

First seen on Joblaze 2 days ago

Last verified on the company career page 16 hours ago

What you'll build

  • Build an execution engine for scoring candidate versions
  • Create task set tooling for production logs
  • Build grader infrastructure for evaluations
  • Develop tooling for human reviewer calibration
  • Build leaderboards and reporting systems

Must have

  • Experience building production platforms
  • Strong statistical literacy
  • Experience evaluating LLM or ML systems

Nice to have

  • Product instinct for internal tools
  • Strong collaboration skills

AI in the day-to-day

We may use automated AI tools to evaluate job applications for efficiency and consistency.

Not disclosed in this posting: years of experience, visa sponsorship.

Benefits

Retirement Savings Plans Remote Work Flexible Time Off Health Insurance

Joblaze summary

The AI Evaluation Infrastructure Engineer at Block focuses on developing the systems and tools necessary for high-quality AI evaluations, ensuring that product teams can quickly and reliably assess model performance. Key skills include experience with production platforms, strong statistical knowledge, and familiarity with machine learning systems. This role is suited for engineers with a background in systems engineering and a keen interest in AI evaluation, particularly those who thrive in collaborative environments. Block's emphasis on innovation and a distributed team culture supports the engineer's impact on AI product development.

Joblaze insights

  • Listed 2 days ago — first seen on Joblaze October 2, 2026. Last confirmed on Block's careers page October 3, 2026.
  • Salary band is above the typical range for AI/ML roles (median ~$190,000).
  • Starts above 94% of 87 comparable mid ai/ml roles in United States that list AI/ML we track (median $152,000 across 38 companies). See AI/ML salary trends
  • AI/ML appears in 51.6% of 465 comparable mid ai/ml roles in United States; distributed systems appears in 1.7% of 465 comparable mid ai/ml roles in United States.

Quick facts

Is the AI Evaluation Infrastructure Engineer role remote?
It's hybrid — Block expects some on-site time in Bay Area, CA, United States.
What's the salary range?
Block lists $263,600–$395,400 for this role.
Where is the role based?
Block is hiring for this position in Bay Area, CA, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, Data Pipelines, distributed systems, statistical analysis.
What seniority level is this role?
Block targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this AI Evaluation Infrastructure Engineer role at Block.

From the original posting

It all started with an idea at Block in 2013. Initially built to take the pain out of peer-to-peer payments, Cash App has gone from a simple product with a single purpose to a dynamic ecosystem, developing unique financial products, including Afterpay/Clearpay, to provide a better way to send, spend, invest, borrow and save to our 50+ million monthly active customers. We want to redefine the world’s relationship with money to make it more relatable, instantly available, and universally accessible.

Today, Cash App has thousands of employees working globally across office and remote locations, with a culture geared toward innovation, collaboration and impact. We’ve been a distributed team since day one, and many of our roles can be done remotely from the countries where Cash App operates. No matter the location, we tailor our experience to ensure our employees are creative, productive, and happy.

The Role

We build AI products, and the quality of our evaluations sets the ceiling for how good those products can be. The speed of our evaluations determines how quickly we can improve them.

We are looking for an engineer to build the infrastructure and tooling that make high-quality AI evaluation possible at Block's scale. You will help teams understand whether a model or product change is actually better, whether a result is statistically meaningful, and whether offline evaluation is predicting what happens with real users.

Our evaluation approach combines offline evals that encode our definition of a good response, online evals that show how people actually respond, and a feedback loop that keeps the two converging. Your work will turn that approach into systems that product teams can use quickly, reliably, and with confidence.

This is a high-impact, early-stage area with broad surface area. You will help decide what to build first, then build the platform that helps teams ship better AI products faster.

You Will

  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.
  • Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.
  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.
  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.
  • Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance, so teams can distinguish real improvements from noise.
  • Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.
  • Build the loop that compares offline scores with online outcomes, identifies eval sets that have stopped predicting reality, and helps teams improve them.
  • Partner with product, engineering, data, and ML teams to make evaluation workflows fast enough and trustworthy enough to become part of everyday development.

You Have

  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.
  • Strong statistical literacy, including comfort with confidence intervals, variance, power, and multiple comparisons.
  • The judgment to identify when a result is meaningful and when it is noise.
  • Experience evaluating LLM or ML systems, or deep systems engineering experience with a strong interest in AI evaluation.
  • Product instinct for internal tools. You understand that leaderboards, annotation tools, and workflows only matter if teams actually use them.
  • A bias toward building reliable, observable systems that other engineers can trust.
  • Strong collaboration skills and the ability to work across ambiguous product, data, and engineering problems.

Success in your first year

  • Product teams can stand up credible evals for new AI surfaces in days.
  • Eval results gate CI and run quickly enough that engineers do not route around them.
  • Judge-to-human agreement is measured, published, and monitored for drift.
  • Launch decisions are not made on results that are within statistical noise.
  • At least one case is documented where online reality disagreed with offline scores, and the evaluation was improved as a result.

Why this matters

AI product development moves quickly, but speed only helps when teams can trust the signal they are using to make decisions. This role will build the systems that make those signals faster, more accurate, and more actionable. Your work will directly influence how Block evaluates, improves, and ships AI products.

Block takes a market-based approach to pay, and pay may vary depending on your location. U.S. locations are categorized into one of four zones based on a cost of labor index for that geographic area. The successful candidate’s starting pay will be determined based on job-related skills, experience, qualifications, work location, and market conditions. These ranges may be modified in the future.

To find a location’s zone designation, please refer to this resource. If a location of interest is not listed, please speak with a recruiter for additional information.

Zone A:
$263,600—$395,400 USD
Zone B:
$263,600—$395,400 USD
Zone C:
$263,600—$395,400 USD
Zone D:
$263,600—$395,400 USD

Application Guidelines

Privacy Policy

Standard company text repeated across Block's postings is omitted here.

Similar positions

Block
Senior Data Engineer, Product
Block · Bay Area, CA, United States
Block
Software Engineer, Data Platform
Block · Bay Area, CA, United States of America
Block
Staff Android Software Engineer, Cash App Consumer Platform
Block · Bay Area, CA, United States of America
Block
Staff Android Software Engineer, Cash App Consumer Platform
Block · Seattle, WA, United States of America
Block
Staff Android Software Engineer, Cash App Consumer Platform
Block · New York, NY, United States of America