Prompt Caching
Cut Claude cost and latency by up to 90% with prompt caching: cache the stable prefix (system prompt, tools, documents, history), pay full price once, then read at ~10%.
Last updated
After this section you can
- Place cache breakpoints so the stable prefix is cached and the volatile tail is not
- Work out write-versus-read costs per model and choose a TTL
- Find silent invalidators and prove caching works from the usage fields
Prompt Caching: Pay Once for the Prefix
Every call resends tools, system prompt and history, and you pay to process them again. Prompt caching lets later calls read that repeated prefix at a fraction of the price and with less wait, as long as its bytes do not change.
Mark where the stable part of your prompt ends with cache_control. The first call writes that prefix at a small premium; calls in the next few minutes that start with exactly the same bytes read it at about a tenth of the price. Change one byte and everything after it is a miss.