1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Grading Agents

Evaluate AI agents in production: task completion metrics, trajectory analysis, and automated agent quality benchmarks.

Last updated

Production11 min readFirst readThe Eval Stack

After this section you can

  • Grade an agent on its end state and its trajectory, and say which failures each grade catches
  • Compute and explain pass@k and pass^k, and choose the right one for a given product
  • Build a harness that runs each task k times in isolation and carries the checks into production
42

Grading Agents: Outcome & Trajectory

An agent does not produce one answer to score. It takes a path through tools and state, a different one each run, and can land on the right result by luck. This is how you grade that.

Key idea

Grade what the agent changed in the world, then check the path it took, and run every task several times. For anything customer-facing, the number that matters is how often it succeeds every time: pass^k.

An agent eval grades the world the agent left behind, the path it took, and how often it does it again
ONE TASK THROUGH THE HARNESS 1 · Task prompt + seeded data + simulated user goal state written first 2 · Agent runs, k times fresh sandbox per trial full trajectory logged trial 1 … k 3a · End state data matches the goal? 3b · Trajectory forbidden calls, steps, loops 3c · Judge only what code cannot check 4 · Report pass@k and pass^k cost and steps per task per tag, with intervals Grade the end state first and the path second, and run every task k times. One run of an agent is an anecdote.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium