1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Metrics & Dashboards

Observability metrics for LLM applications: latency, token usage, cost tracking, and quality scoring dashboards.

Last updated

Production9 min readFirst readGrading Agents

After this section you can

  • Compute cost per resolved task and explain why it beats cost per call
  • Define the quality, latency (TTFT and task duration), tool-error and escalation metrics that drive decisions
  • Lay out a dashboard where every tile has a threshold, an owner and a next action
43

Metrics & Dashboards: What to Track

Evals tell you whether a version is good. Production metrics tell you whether the live agent is still good today, and which team should act. Five numbers cover most of it.

Key idea

A metric earns a place on the dashboard only if someone would do something different when it moves. Count per task, split by segment, and give each number a threshold, an owner and a next action.

Five numbers, each with a threshold and a next action. Anything else is a drill-down.
SUPPORT AGENT · LAST 7 DAYS QUALITY Resolution rate 80% goal reached · 7 days alert: under 74% → read 20 failed traces LATENCY TTFT p95 · task p95 1.4 · 11 s first token · whole task alert: task p95 over 15 s → find the slow span COST Cost / resolved task $0.47 spend ÷ resolved tasks alert: 1.5× baseline → split by route, model TOOLS Tool error rate 2.1% is_error results ÷ calls alert: any tool over 5% → page that tool’s owner HANDOFFS Escalation rate 14% tasks handed to a human alert: +3 pts week on week → sample 30 by reason Every tile splits by task type · tenant · model version, with deploy markers on the time axis Numbers are illustrative. Set each threshold from your own baseline, and give every tile an owner.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium