Grading Agents
Evaluate AI agents in production: task completion metrics, trajectory analysis, and automated agent quality benchmarks.
Last updated
After this section you can
- Grade an agent on its end state and its trajectory, and say which failures each grade catches
- Compute and explain pass@k and pass^k, and choose the right one for a given product
- Build a harness that runs each task k times in isolation and carries the checks into production
Grading Agents: Outcome & Trajectory
An agent does not produce one answer to score. It takes a path through tools and state, a different one each run, and can land on the right result by luck. This is how you grade that.
Grade what the agent changed in the world, then check the path it took, and run every task several times. For anything customer-facing, the number that matters is how often it succeeds every time: pass^k.