← Back to results

Research Engineer

Join Ando as a Research Engineer to work on AI agents in a messaging platform, focusing on real data and impactful research.

Location
San Francisco, United States
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 2 months ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Ando → Save job Scanned from ando.so

What you'll build

  • Build evaluation from real data
  • Run experiments that ship
  • Decide what we build versus who we partner with
  • Publish when we have something real

Must have

  • Strong applied research background
  • Depth in model evaluation, benchmarking, and/or failure analysis
  • Strong technical communication

Nice to have

  • Familiarity with simulation
  • Human-in-the-loop evaluation
  • Memory for multi-party long-running settings

AI in the day-to-day

AI agents take on work alongside human teammates in a messaging platform.

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Benefits

Gym Membership Equity/Stock Options Health Insurance

Joblaze summary

In this role, the Research Engineer at Ando will focus on developing evaluation benchmarks and conducting experiments that directly influence the company's messaging platform. Key skills include applied research expertise in model evaluation and the ability to handle real-world data complexities, with an emphasis on technical communication. This position is ideal for someone with a strong background in research and a hands-on approach to problem-solving, particularly in environments where data is messy and nuanced. The team is small and collaborative, emphasizing practical applications of research to enhance product quality.

Joblaze insights

  • Listed yesterday — first seen on Joblaze September 28, 2026. Last confirmed on Ando's careers page September 28, 2026.
  • AI/ML appears in 52.3% of 459 comparable mid ai/ml roles in United States.

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, benchmarking, evaluation, failure analysis, labeling infrastructure.
What seniority level is this role?
Ando targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Research Engineer role at Ando.

From the original posting

Ando is a messaging platform where AI agents take on work alongside their human teammates. We’re rebuilding Slack from the ground up around two core ideas: durable memory and agents as first-class participants.

We have real, longitudinal, multi-party workspace data, and human / agent users whose behavior tells you whether the systems you're designing are actually meaningfully improving. Working with live business communication means permissions, redaction, and security are things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.

Some concrete research problems on our plate right now:

  • Proactivity: An agent embedded in a team's channels has to decide, message by message, whether to ignore, quietly track, or intervene. Evaluating that judgment means building benchmarks where the ground truth includes silence. Existing agent benchmarks are almost entirely reactive.

  • Memory: What should a workspace agent remember across weeks and months of participation, in what representation, and how do you measure whether memory is helping versus hurting agent usefulness?

  • Continual Learning: How much does basic memory, retrieval, and context impact agent performance, vs where do we genuinely need RL and continual learning?

What you’ll be doing at Ando

You'll be one of the first members of a small research team, working close to both the data and the product.

  • Build evaluation from real data - Mine production workspace data (carefully, with consent and redaction pipelines you'll help design) into benchmarks and labeled datasets. Design label schemas, run labeling with real inter-rater rigor, and build the harnesses that make expert judgment cheap to capture and hard to corrupt.

  • Run experiments that ship - Initial work happens on offline workspace data; the destination is production systems used by every Ando customer. The distance between "the benchmark improved" and "the feature shipped" should be weeks and you'll own both ends.

  • Decide what we build versus who we partner with - We won't do everything in-house. Part of the job is evaluating frontier vendors and research teams (eval infrastructure, observability, continual-learning tooling) and choosing who we build with.

  • Publish when we have something real - We expect the team to publish as we make meaningful progress. But ideas are not Ando's moat; execution, product quality, and customer experience are. Research here is in service of customers first, and the publications will be better for it: they'll describe things that actually worked on real data.

What we're looking for

  • Strong applied research background, with depth in model evaluation, benchmarking, and/or failure analysis. You've built evals you trusted enough to make decisions with.

  • Evidence over credentials. Work samples or code that demonstrate the skills: eval frameworks, benchmark suites, failure-analysis reports or tooling, labeling infrastructure. Show us something you built to find out whether a system actually worked.

  • Strong technical communication. You can explain complex ideas simply and hold high-bandwidth, generative technical conversations with researchers and with our product team.

  • Comfort with mess. Real workspace data is incomplete, ambiguous, and full of edge cases that break clean abstractions. You treat that as signal, not noise.

  • Bonus: familiarity with simulation (Park et al.), human-in-the-loop evaluation (Scale HIL leaderboard), Cartridges and related context/memory-compression work, memory for multi-party long-running settings, or agent observability standards (setting up Langsmith or similar).

Hiring process

  1. 30 min intro call

  2. Technical conversation - walk us through an app you've shipped

  3. Paid take-home or IRL work trial

Benefits

  • Free Equinox membership & other health perks

  • Generous equity grant vested over 4 years

  • Health, dental, vision insurance

Similar positions

Ando
Product Engineer
Ando · San Francisco, United States
Ando
Product Designer
Ando · San Francisco, United States
Cognition
Research Engineer, Post-Training
Cognition · San Francisco
Cognition
Research Engineer, Mid-Training
Cognition · San Francisco
WisdomAI
Research Engineer/ Applied Scientist
WisdomAI · San Mateo, United States