✓Design a four-layer eval stack from a CI gate through golden sets and A/B tests to production monitoring
✓Build a golden set and an LLM judge that survive scrutiny: code graders first, calibrated and bias-checked judging
✓Instrument an agent with OpenTelemetry GenAI spans so every layer reads the same traces
41
The Eval Stack: Offline to Online
A score on a test set is one layer of evidence. This is the whole stack that lets you change a prompt, a tool or a model on a Tuesday and know by Friday whether users are better off.
Key idea
Evaluation is a pipeline of gates. Cheap, fast checks run on every change; slower, truer ones run on live traffic; production failures flow back as new test cases. Traces are the shared record every gate reads.
Four gates between a change and your users, all reading the same traces