1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Context Management

Context editing, compaction and memory as Claude API features: what each does to the conversation, and why the compaction block must be appended back verbatim.

Last updated

Production11 min readFirst readPrompt CachingAdaptive Thinking & Effort

After this section you can

  • Choose between context editing, compaction and memory for a long-running agent
  • Wire up compaction without losing its block, and read its cost from usage.iterations
  • Explain why trimming history yourself breaks the newest models, and what replaces it
32

Context Management: Clear, Compact, Persist

Three API features with three jobs: clearing finished material, summarising a conversation that will not fit, and keeping what must outlive it. On the newest models they are also the only safe way to shrink history.

Key idea

Long agent runs fill the window with material that has done its job. Context editing clears it, compaction summarises it, memory writes it somewhere that outlives the session. The first two run on the server, so your history stays append-only, which the newest models now require.

Same full window, three different things done to it
BEFORE AFTER CONTEXT EDITING clear_tool_uses_20250919 COMPACTION compact_20260112 MEMORY memory_20250818 sys tool result result sys cleared cleared sys sys compaction room again sys write sys read session 2 starts here session 1, all of it Every turn stays. Old tool results become a placeholder on the server; your copy is untouched. Earlier turns become one server-written block. Append the full response.content or it is lost. Only what the model wrote to its memory directory reaches the next session.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium