Adaptive Thinking & Effort
Let Claude reason before it answers: adaptive thinking, the five effort levels, why budget_tokens is gone, summarized vs omitted thinking blocks, and preserving them across tool calls.
Last updated
After this section you can
- Configure adaptive thinking and effort per model, and size max_tokens for thinking plus the answer
- Explain what replaced budget_tokens, what display controls and what it costs
- Carry thinking blocks correctly through a tool loop and choose when to raise or lower effort
Adaptive Thinking & Effort: Reasoning on Demand
Current Claude models reason before they answer, and decide for themselves how much. Your controls are one effort setting, how the reasoning is displayed, and passing it back correctly in tool loops.
Turn thinking on with {"type": "adaptive"} and steer it with output_config.effort. There is no token budget any more: Claude picks the depth per turn, effort bounds it, and you pay for every thinking token whether you see the text or not.