← Back to results

Product engineer, Agent

Build product experiences for AI agents, owning problems end-to-end in a fast-paced environment.

Location
San Francisco
Compensation
Not disclosed
Level
mid
Type
full time

Posted by employer 4 weeks ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Judgment Labs → Save job Scanned from judgmentlabs.ai

Skills & Technologies

AI in the day-to-day

Judgment is the learning infrastructure for AI agents, improving them through experience and structured signals.

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Joblaze summary

In this role, the Product Engineer at Judgment Labs focuses on developing and refining the infrastructure that enhances AI agents through real-world experience and feedback. Key skills include building scalable production systems and a strong grasp of LLMs or agent technologies, alongside effective communication with customers to address their needs. This position is ideal for someone with a proactive, problem-solving mindset who thrives in dynamic environments and enjoys taking ownership of projects from conception to execution.

Joblaze insights

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, LLMs.
What seniority level is this role?
Judgment Labs targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Product engineer, Agent role at Judgment Labs.

From the original posting

Product Engineer - Agents Job Description

The Role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:

  1. We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

  2. Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals

  3. Teams close the loop, shipping agent improvements validated against real production evidence

You'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.

What You Will Accomplish

  • Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.

  • Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.

  • Agent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the "program" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes?

  • Swarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports.

  • The improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.

  • The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

  • Judgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development.

What You'll Bring

  • Experience building and scaling end-to-end production systems, from data layer to UI

  • Strong technical problem-solving skills, especially in fast-changing, ambiguous environments

  • A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship

  • Hands-on experience building with LLMs or agents, or the drive to get there fast

  • Comfort working directly with customers to understand their needs and solve real-world problems

  • Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences

Similar positions

Judgment Labs
Product engineer, full stack
Judgment Labs · San Francisco
Judgment Labs
Applied AI Engineer
Judgment Labs · San Francisco
Judgment Labs
Backend/Infra Engineer
Judgment Labs · San Francisco
Fieldguide
Senior AI Engineer
Fieldguide · San Francisco, CA
ChipAgents
Senior Design Engineer
ChipAgents · San Jose