1:1 mentoring with Big Tech AI engineers
LLM & AgenticFree

The Agent Loop

Build a complete tool-calling AI agent in 15 lines of Python. Understand the core agent loop pattern that powers all LLM agents.

Last updated

After this section you can

  • Write a working tool-calling agent loop from memory and handle every stop_reason on purpose
  • Explain why state is the messages array and estimate how its input tokens grow per step
  • Add the guards production needs: step and cost ceilings, timeouts, capped results and errors returned as results
08

The Agent Loop: Think → Act → Observe

Strip away the frameworks and every agent is the same fifteen lines: a model, some tools and a bounded loop. Learn where its state lives, what controls it and how it fails, and every agent product becomes legible.

Key idea

An agent is a model calling tools in a loop until it decides it is done, or until one of your guards stops it. The state is the messages array. The control signal is stop_reason. Everything else is decoration.

Every agent is this loop: call the model, branch on stop_reason, run tools, append, repeat
1 · messages[] the only state task + every turn so far 2 · model call messages.create( tools, messages) 3 · stop_reason? the only control signal read it before anything else done return the text blocks and stop 4 · your code runs tools failures become is_error results your guards max steps · budget · timeout end_turn tool_use 5 · append assistant turn + all results State lives in messages[]. Control lives in stop_reason. The stops you cannot skip are the ones you write.

The loop in fifteen lines

This is a complete agent. The model thinks and picks a tool; your code acts; the result is appended and the model observes it on the next call. Two details carry the contract from Inside a Tool Call: append the model’s full content, and send every result back in one user message.

def run_agent(task: str, max_steps: int = 20) -> str:
    messages = [{"role": "user", "content": task}]
    for _ in range(max_steps):                         # a budget, never while True
        resp = client.messages.create(
            model="claude-opus-5", max_tokens=16000,
            tools=TOOLS, messages=messages)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason == "end_turn":
            return "".join(b.text for b in resp.content if b.type == "text")
        if resp.stop_reason != "tool_use":             # max_tokens, pause_turn, refusal
            raise AgentStopped(resp.stop_reason, messages)
        results = [run_tool(b) for b in resp.content if b.type == "tool_use"]
        messages.append({"role": "user", "content": results})   # one message
    raise AgentStopped("max_steps", messages)

The answer is the joined text blocks, not content[0].text: with thinking on by default on current models, the first block is often a thinking block.

You rarely need to write this by hand. The SDK’s Tool Runner (client.beta.messages.tool_runner) runs the same loop for tools you define, with max_iterations as the step guard. When to use it, the Agent SDK or a hosted harness instead is in Who Owns the Loop.

State is messages[], and it only grows

The loop remembers nothing except the array it re-sends. After one cycle it holds four turns, and every later call re-reads all of them.

messages[] after one tool call
user

Is order A-1042 going to be late?

assistant
[{"type": "text", "text": "I'll check the order first."},
 {"type": "tool_use", "id": "toolu_01", "name": "get_order",
  "input": {"order_id": "A-1042"}}]
stop_reason: "tool_use". Appended whole, text and tool_use together.
user
[{"type": "tool_result", "tool_use_id": "toolu_01",
  "content": "{\"status\": \"shipped\", \"eta\": \"2026-09-29\"}"}]
Results travel in a user turn, keyed by tool_use_id.
assistant

It shipped and is due on 29 September, so it is on time.

stop_reason: "end_turn". The loop returns.

Growth is the cost driver. Take 3,000 tokens of system prompt and tools, and 2,000 tokens added per step (the call plus its result). Call n reads 3,000 + 2,000 × (n − 1) tokens, so a 20-step run looks like this:

41Kinput tokens read by call 20 alone
440Kinput tokens across all 20 calls
$2.20of uncached input at $5 per million

Most of each call is a prefix the previous call already sent, which is why Prompt Caching changes an agent’s bill so much. When the history itself gets too long, the newest Claude models expect history to stay append-only: use server-side compaction or context editing rather than deleting old turns (see Context Management).

Control is stop_reason

The loop has one branch point. Handle each value explicitly; treating “anything but tool_use” as success ships truncated and refused answers as if they were finished.

stop_reasonWhat it meansWhat the loop does
tool_usethe model wants one or more tools runrun every tool_use block, append all results, call again
end_turnthe model thinks it is donereturn the text (and verify the outcome if you can)
max_tokensthe response was cut offnever run a tool call from it; raise max_tokens (stream large values) or stop
pause_turna server-tool loop pausedappend the assistant turn unchanged and call again, with no “Continue” message
refusalthe model declinedstop and surface it; check stop_details, do not retry blindly

The full list, including stop_sequence, is in Messages API. The Tool Runner does not auto-resume pause_turn, so handle it if you use server tools.

Tool errors go back as results

If a tool raises inside your loop, the request is left with an unanswered tool_use and the model never learns what happened. Catch it and send it back. Models read TimeoutError: inventory API after 5s and retry, switch tools or tell the user.

def run_tool(block) -> dict:
    try:
        out = str(TOOL_FNS[block.name](**block.input))
        if len(out) > 8000:
            out = out[:8000] + "\n[truncated: 8,000 of %d chars]" % len(out)
        return {"type": "tool_result", "tool_use_id": block.id, "content": out}
    except Exception as e:                  # unknown tool, bad args, timeout
        return {"type": "tool_result", "tool_use_id": block.id,
                "content": f"{type(e).__name__}: {e}", "is_error": True}
Raise crashes the run
  • the next request fails: a tool_use has no result
  • the model cannot adapt
  • one flaky API ends a 30-step task
Return is_error: true
  • every call gets a result
  • the error text tells the model what to do next
  • your guards still bound the retries

Guards: the stops that belong to you

The model decides when it is done. You decide when it has had enough. Every guard ends the same way: stop, record why, and return partial progress marked as incomplete, never as success.

GuardWhat it boundsA starting point
Max stepsloops and thrash20 for a support task, more for coding; set from real traces
Token or cost ceilingrunaway spendsum usage per call and stop at the ceiling; a task_budget also lets the model pace itself (see Agent Cost Control)
Wall-clock timeouthung tools, slow runsone per tool call, one for the whole run
Result size capone result flooding contexttruncate and say so, as in run_tool above
Repeat detectorthe same call over and overhash (tool, arguments); on the third repeat, stop or tell the model
Approval gateirreversible side effectspause before refunds, emails, deletes (see Approval Gates)

A long run can also die for reasons outside the loop: a deploy, a crash, an expired container. Surviving that is Agent Durability.

How the loop fails

FailureWhat you see in the traceFix
Infinite loopsame tool, same arguments, step after steprepeat detector, max steps
Thrashmany different tools, no progressstep budget; clearer tool descriptions (see Tool Surface Design)
Premature finishend_turn saying “done” when nothing was donecheck the outcome in code (tests, a lookup) before returning
Silent tool failurethe model reasons over an empty or garbage resultreturn failures with is_error and actionable text
Context growtheach call slower and pricier; early instructions followed lesscap results; compaction or context editing
Runaway costa $40 run for a $0.04 questiona cost ceiling in the loop, alerts per run
Common mistakes

Appending only the answer text instead of the full content (tool calls and thinking blocks go missing and the next request fails). Sending each tool_result in its own user message (the model learns to stop calling tools in parallel). Writing while True with no step or cost guard.

Interview angle

“Build me an agent on a whiteboard.” Write the fifteen lines first and narrate them: messages is the state, stop_reason is the control, tools run in my code. Then add the guards out loud: step cap, cost ceiling, timeouts, errors returned as results, outputs capped, approvals on anything irreversible. Framework names are optional; the loop is not.

Check yourself
Your loop returns the first content block’s text whenever stop_reason is anything other than tool_use. Name two ways this breaks on a current model.

With thinking on, content[0] is usually a thinking block, so there is no .text to return. And max_tokens or refusal are returned as if they were finished answers. Branch on each stop reason and join the text blocks.

A tool times out and your code raises. The run ends with an API error on the next request. Why, and what is the fix?

The assistant turn contained a tool_use that never got a tool_result, which the API rejects. Catch the exception and return it as a result with is_error: true, so the model can retry or change course.

A 30-step research run costs ten times what the first few calls suggested. What is happening, and which two levers do you pull first?

Each call re-reads the whole growing history, so input tokens grow roughly with the square of the step count. Cache the stable prefix, and cap tool result sizes; then use compaction or context editing if the history still grows too long.

Recap
  • An agent is a bounded loop: call the model, branch on stop_reason, run tools, append, repeat.
  • messages is the only state. Append the full content, and all results in one user message.
  • Handle every stop reason on purpose; only end_turn means done.
  • Tool failures go back as is_error results. Guards (steps, cost, time, size, repeats) are yours.
  • History grows every step: cache the prefix, and use compaction or context editing rather than deleting turns.

Next: The Agentic Spectrum: Workflows to Autonomy, and why most problems need less than a full agent.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium