1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Agent Cost Control

Control what an agent costs: effort levels, task budgets vs session budgets vs max_tokens, prompt-cache economics, model routing, and per-turn token accounting.

Last updated

After this section you can

  • Split an agent’s bill into cache reads, cache writes, uncached input and output before choosing a fix
  • Use effort as the first cost dial, and choose between max_tokens, task budgets and hard spend caps
  • Route work across models without breaking the cache, and keep a per-turn cost ledger
33

Agent Cost Control: Effort, Budgets, Caching

An agent’s bill is a per-turn choice about depth, a prefix you either reuse or pay for again, a price per token set by the model, and ceilings that stop a run in different ways.

Key idea

Split the bill before you fix it. On a long run, most tokens are input the model re-reads every turn, so the cache sets most of the price. Effort is the first dial that trades quality for cost, routing changes the price per token, and the ceilings only decide where a run stops.

Where one agent run’s money goes, and which lever moves each part
ONE 40-TURN RUN ON CLAUDE OPUS 5 · ABOUT $3.75 INPUT · 2.4M tokens re-read over 40 turns OUTPUT · 60K tokens cache reads $1.10 writes $0.75 $0.40 output + thinking $1.50 2.2M × $0.50/M 120K × $6.25/M 80K × $5/M 60K × $25/M Caching: a stable prefix keeps 2.2M of 2.4M input tokens at 0.1× Effort: thinking depth and number of turns Routing: the model sets every price above (Sonnet 5: $2 / $10 per million, Haiku 4.5: $1 / $5) Ceilings (max_tokens, task budget, session budget) decide where a run stops, not what each token costs.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium