1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Adaptive Thinking & Effort

Let Claude reason before it answers: adaptive thinking, the five effort levels, why budget_tokens is gone, summarized vs omitted thinking blocks, and preserving them across tool calls.

Last updated

Production10 min readFirst readMessages APITool Use with Claude

After this section you can

  • Configure adaptive thinking and effort per model, and size max_tokens for thinking plus the answer
  • Explain what replaced budget_tokens, what display controls and what it costs
  • Carry thinking blocks correctly through a tool loop and choose when to raise or lower effort
24

Adaptive Thinking & Effort: Reasoning on Demand

Current Claude models reason before they answer, and decide for themselves how much. Your controls are one effort setting, how the reasoning is displayed, and passing it back correctly in tool loops.

Key idea

Turn thinking on with {"type": "adaptive"} and steer it with output_config.effort. There is no token budget any more: Claude picks the depth per turn, effort bounds it, and you pay for every thinking token whether you see the text or not.

You set how far it may think. Claude decides how far it does.
Request thinking: adaptive effort: low … max max_tokens caps thinking and answer together Claude picks the depth lookup bug hunt hard proof effort sets how far it can go Response content thinking text empty by default · signature text or tool_use the answer, or the next action 1 · 2 · 3 · next turn: send the thinking blocks back exactly as received Thinking tokens are output tokens: billed, and counted against max_tokens, even when their text is hidden.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium