1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Prompt Caching

Cut Claude cost and latency by up to 90% with prompt caching: cache the stable prefix (system prompt, tools, documents, history), pay full price once, then read at ~10%.

Last updated

Production11 min readFirst readMessages APITokens, Context Windows & Sampling

After this section you can

  • Place cache breakpoints so the stable prefix is cached and the volatile tail is not
  • Work out write-versus-read costs per model and choose a TTL
  • Find silent invalidators and prove caching works from the usage fields
25

Prompt Caching: Pay Once for the Prefix

Every call resends tools, system prompt and history, and you pay to process them again. Prompt caching lets later calls read that repeated prefix at a fraction of the price and with less wait, as long as its bytes do not change.

Key idea

Mark where the stable part of your prompt ends with cache_control. The first call writes that prefix at a small premium; calls in the next few minutes that start with exactly the same bytes read it at about a tenth of the price. Change one byte and everything after it is a miss.

The cache is a prefix: one changed byte invalidates everything after it
RENDER ORDER IS CACHE ORDER tools fixed, sorted by name system frozen: no dates, no IDs messages: earlier turns append-only, never edited new user turn different every call breakpoint 1 breakpoint 2 NEXT CALL · SAME PREFIX read from cache: about 0.1× input price full price NEXT CALL · A TIMESTAMP ADDED TO THE SYSTEM PROMPT still read changed at the system prompt: everything from here is written again at 1.25× Stable content first, volatile content last, breakpoints on the boundaries between them.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium