← Back to results

Applied AI Researcher, Benchmarking

Join Distyl AI as an Applied AI Researcher to redefine software usage and drive innovative benchmarking in AI systems.

Location
San Francisco
Compensation
$150k–$250k/yr
Level
senior
Type
full time · Hybrid

Posted by employer 11 months ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

AI in the day-to-day

You should be using tools like ChatGPT, Cursor, and Perplexity to accelerate your workflow.

Not disclosed in this posting: years of experience, visa sponsorship.

Benefits

401k Match Wellness Benefits Equity/Stock Options Complimentary Lunches Flexible Time Off Health Insurance

Joblaze summary

In the role of Applied AI Researcher focused on Benchmarking at Distyl AI, the individual will design and implement evaluation frameworks that assess the performance and reliability of AI systems in real-world scenarios. Candidates should possess strong analytical skills and experience in creating benchmarks, alongside a solid understanding of advanced AI techniques and programming. This position is ideal for researchers with a proven track record in AI who thrive in a dynamic environment that values innovative thinking and practical applications. Distyl AI fosters a mission-driven culture, emphasizing collaboration and the impact of AI on enterprise operations.

Joblaze insights

Quick facts

Is the Applied AI Researcher, Benchmarking role remote?
It's hybrid — Distyl AI expects some on-site time in San Francisco.
What's the salary range?
Distyl AI lists $150,000–$250,000 for this role.
Where is the role based?
Distyl AI is hiring for this position in San Francisco.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, ChatGPT, Cursor, Perplexity.
What seniority level is this role?
Distyl AI targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Applied AI Researcher, Benchmarking role at Distyl AI.

From the original posting

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.

What We Are Looking For

At Distyl we’re pushing the envelope of AI utilization in enterprise. This requires creative researchers who don’t just want to drive incremental improvements on benchmarks or optimize an existing process but instead are looking to creatively redefine how software is used.

Our researchers come from many academic backgrounds but have strong research track records, operate in an AI-native way, and would be bored staying on the rails of a traditional research org.

Key Responsibilities

  • The Benchmarking team defines how progress is measured. Researchers design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact. They construct benchmarks that reflect real-world complexity. Their systems become the standard by which new architectures, techniques, and releases are judged.

  • Researchers in Benchmarking explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment. They investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability. Their insights drive both Distyl’s internal research priorities and industry-wide standards.

Who You Are

  • Experience Designing and Running Evaluations: You’ve built or maintained benchmarks, test suites, or experimental frameworks to measure model or system performance

  • Statistical and Analytical Rigor: You design fair, reproducible experiments and can extract signal from noisy empirical results

  • Experience Building with Models, Not Just Building Models: We develop intelligent systems using models rather than training or fine-tuning them. Ideal candidates have expertise in compound AI systems, agentic collaboration, and associated techniques (ensembling, ReAct, graph-of-thoughts, etc.)

  • Proven Track Record of Research Results: Whether you’ve published in top journals, posted amazing work on twitter, or somewhere else we want to see what you've done

  • Uses AI Every Day: Before you can revolutionize someone else’s workflow, you need to revolutionize yours. You should be using tools like ChatGPT, Cursor, and Perplexity to accelerate your workflow

  • Strong Programming and Data Analysis Skills: While you might not consider yourself a software engineer you need to be able to build prototypes of your ideas and then perform the experiments to prove the effectiveness to a F500 Head of AI

  • Biases Towards Showing vs Telling: Our customers want to see the power of AI today vs discuss the most elegant idea that will take 5 years to realize

What We Offer

  • The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot

  • Complimentary in-office lunches and snacks provided

  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems

  • Ownership of high-impact projects across top enterprises

  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office.

#LI-Hybrid

We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Similar positions

Distyl AI
Applied AI Researcher, AI Systems
Distyl AI · San Francisco
Distyl AI
Distyl AI
Distyl AI
Research Engineers, Data
Distyl AI · San Francisco
Distyl AI
Applied AI Researcher, Post-Training
Distyl AI · San Francisco