Agent Cost Control
Control what an agent costs: effort levels, task budgets vs session budgets vs max_tokens, prompt-cache economics, model routing, and per-turn token accounting.
Last updated
After this section you can
- Split an agent’s bill into cache reads, cache writes, uncached input and output before choosing a fix
- Use effort as the first cost dial, and choose between max_tokens, task budgets and hard spend caps
- Route work across models without breaking the cache, and keep a per-turn cost ledger
Agent Cost Control: Effort, Budgets, Caching
An agent’s bill is a per-turn choice about depth, a prefix you either reuse or pay for again, a price per token set by the model, and ceilings that stop a run in different ways.
Split the bill before you fix it. On a long run, most tokens are input the model re-reads every turn, so the cache sets most of the price. Effort is the first dial that trades quality for cost, routing changes the price per token, and the ceilings only decide where a run stops.