The Agent Loop
Build a complete tool-calling AI agent in 15 lines of Python. Understand the core agent loop pattern that powers all LLM agents.
Last updated
After this section you can
- Write a working tool-calling agent loop from memory and handle every stop_reason on purpose
- Explain why state is the messages array and estimate how its input tokens grow per step
- Add the guards production needs: step and cost ceilings, timeouts, capped results and errors returned as results
The Agent Loop: Think → Act → Observe
Strip away the frameworks and every agent is the same fifteen lines: a model, some tools and a bounded loop. Learn where its state lives, what controls it and how it fails, and every agent product becomes legible.
An agent is a model calling tools in a loop until it decides it is done, or until one of your guards stops it. The state is the messages array. The control signal is stop_reason. Everything else is decoration.
The loop in fifteen lines
This is a complete agent. The model thinks and picks a tool; your code acts; the result is appended and the model observes it on the next call. Two details carry the contract from Inside a Tool Call: append the model’s full content, and send every result back in one user message.
def run_agent(task: str, max_steps: int = 20) -> str:
messages = [{"role": "user", "content": task}]
for _ in range(max_steps): # a budget, never while True
resp = client.messages.create(
model="claude-opus-5", max_tokens=16000,
tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "end_turn":
return "".join(b.text for b in resp.content if b.type == "text")
if resp.stop_reason != "tool_use": # max_tokens, pause_turn, refusal
raise AgentStopped(resp.stop_reason, messages)
results = [run_tool(b) for b in resp.content if b.type == "tool_use"]
messages.append({"role": "user", "content": results}) # one message
raise AgentStopped("max_steps", messages)
The answer is the joined text blocks, not content[0].text: with thinking on by default on current models, the first block is often a thinking block.
You rarely need to write this by hand. The SDK’s Tool Runner (client.beta.messages.tool_runner) runs the same loop for tools you define, with max_iterations as the step guard. When to use it, the Agent SDK or a hosted harness instead is in Who Owns the Loop.
State is messages[], and it only grows
The loop remembers nothing except the array it re-sends. After one cycle it holds four turns, and every later call re-reads all of them.
Is order A-1042 going to be late?
[{"type": "text", "text": "I'll check the order first."},
{"type": "tool_use", "id": "toolu_01", "name": "get_order",
"input": {"order_id": "A-1042"}}]
stop_reason: "tool_use". Appended whole, text and tool_use together.[{"type": "tool_result", "tool_use_id": "toolu_01",
"content": "{\"status\": \"shipped\", \"eta\": \"2026-09-29\"}"}]
Results travel in a user turn, keyed by tool_use_id.It shipped and is due on 29 September, so it is on time.
stop_reason: "end_turn". The loop returns.Growth is the cost driver. Take 3,000 tokens of system prompt and tools, and 2,000 tokens added per step (the call plus its result). Call n reads 3,000 + 2,000 × (n − 1) tokens, so a 20-step run looks like this:
Most of each call is a prefix the previous call already sent, which is why Prompt Caching changes an agent’s bill so much. When the history itself gets too long, the newest Claude models expect history to stay append-only: use server-side compaction or context editing rather than deleting old turns (see Context Management).
Control is stop_reason
The loop has one branch point. Handle each value explicitly; treating “anything but tool_use” as success ships truncated and refused answers as if they were finished.
stop_reason | What it means | What the loop does |
|---|---|---|
tool_use | the model wants one or more tools run | run every tool_use block, append all results, call again |
end_turn | the model thinks it is done | return the text (and verify the outcome if you can) |
max_tokens | the response was cut off | never run a tool call from it; raise max_tokens (stream large values) or stop |
pause_turn | a server-tool loop paused | append the assistant turn unchanged and call again, with no “Continue” message |
refusal | the model declined | stop and surface it; check stop_details, do not retry blindly |
The full list, including stop_sequence, is in Messages API. The Tool Runner does not auto-resume pause_turn, so handle it if you use server tools.
Tool errors go back as results
If a tool raises inside your loop, the request is left with an unanswered tool_use and the model never learns what happened. Catch it and send it back. Models read TimeoutError: inventory API after 5s and retry, switch tools or tell the user.
def run_tool(block) -> dict:
try:
out = str(TOOL_FNS[block.name](**block.input))
if len(out) > 8000:
out = out[:8000] + "\n[truncated: 8,000 of %d chars]" % len(out)
return {"type": "tool_result", "tool_use_id": block.id, "content": out}
except Exception as e: # unknown tool, bad args, timeout
return {"type": "tool_result", "tool_use_id": block.id,
"content": f"{type(e).__name__}: {e}", "is_error": True}
- the next request fails: a
tool_usehas no result - the model cannot adapt
- one flaky API ends a 30-step task
is_error: true- every call gets a result
- the error text tells the model what to do next
- your guards still bound the retries
Guards: the stops that belong to you
The model decides when it is done. You decide when it has had enough. Every guard ends the same way: stop, record why, and return partial progress marked as incomplete, never as success.
| Guard | What it bounds | A starting point |
|---|---|---|
| Max steps | loops and thrash | 20 for a support task, more for coding; set from real traces |
| Token or cost ceiling | runaway spend | sum usage per call and stop at the ceiling; a task_budget also lets the model pace itself (see Agent Cost Control) |
| Wall-clock timeout | hung tools, slow runs | one per tool call, one for the whole run |
| Result size cap | one result flooding context | truncate and say so, as in run_tool above |
| Repeat detector | the same call over and over | hash (tool, arguments); on the third repeat, stop or tell the model |
| Approval gate | irreversible side effects | pause before refunds, emails, deletes (see Approval Gates) |
A long run can also die for reasons outside the loop: a deploy, a crash, an expired container. Surviving that is Agent Durability.
How the loop fails
| Failure | What you see in the trace | Fix |
|---|---|---|
| Infinite loop | same tool, same arguments, step after step | repeat detector, max steps |
| Thrash | many different tools, no progress | step budget; clearer tool descriptions (see Tool Surface Design) |
| Premature finish | end_turn saying “done” when nothing was done | check the outcome in code (tests, a lookup) before returning |
| Silent tool failure | the model reasons over an empty or garbage result | return failures with is_error and actionable text |
| Context growth | each call slower and pricier; early instructions followed less | cap results; compaction or context editing |
| Runaway cost | a $40 run for a $0.04 question | a cost ceiling in the loop, alerts per run |
Appending only the answer text instead of the full content (tool calls and thinking blocks go missing and the next request fails). Sending each tool_result in its own user message (the model learns to stop calling tools in parallel). Writing while True with no step or cost guard.
“Build me an agent on a whiteboard.” Write the fifteen lines first and narrate them: messages is the state, stop_reason is the control, tools run in my code. Then add the guards out loud: step cap, cost ceiling, timeouts, errors returned as results, outputs capped, approvals on anything irreversible. Framework names are optional; the loop is not.
Your loop returns the first content block’s text whenever stop_reason is anything other than tool_use. Name two ways this breaks on a current model.
With thinking on, content[0] is usually a thinking block, so there is no .text to return. And max_tokens or refusal are returned as if they were finished answers. Branch on each stop reason and join the text blocks.
A tool times out and your code raises. The run ends with an API error on the next request. Why, and what is the fix?
The assistant turn contained a tool_use that never got a tool_result, which the API rejects. Catch the exception and return it as a result with is_error: true, so the model can retry or change course.
A 30-step research run costs ten times what the first few calls suggested. What is happening, and which two levers do you pull first?
Each call re-reads the whole growing history, so input tokens grow roughly with the square of the step count. Cache the stable prefix, and cap tool result sizes; then use compaction or context editing if the history still grows too long.
- An agent is a bounded loop: call the model, branch on
stop_reason, run tools, append, repeat. messagesis the only state. Append the fullcontent, and all results in one user message.- Handle every stop reason on purpose; only
end_turnmeans done. - Tool failures go back as
is_errorresults. Guards (steps, cost, time, size, repeats) are yours. - History grows every step: cache the prefix, and use compaction or context editing rather than deleting turns.