Agent Durability
Make a long agent run survive a crash: the window where a side effect lands before anything records it, idempotency keys that survive replay, and what to checkpoint.
Last updated
After this section you can
- Locate the window in each loop step where a crash repeats a side effect, and make the replay a no-op
- Derive idempotency keys that survive a replay, and order writes so the keys stay stable
- Checkpoint the transcript append-only and verbatim, and account for cold caches and reset budgets on resume
Agent Durability: Surviving a Crash Mid-Run
A long agent run is a distributed transaction nobody designed as one. When it dies at step 40, the useful question is what step 39 already did to the world.
Replaying a model call costs tokens; replaying a tool can refund a customer twice. Each step has a window between the side effect and the record of it, and outside a single database transaction you cannot close it. So save the transcript append-only, save each model turn before running its tools, and key every side effect so a replay is a no-op.