1:1 mentoring with Big Tech AI engineers
LLM & AgenticFree

ReAct Pattern

ReAct (Reasoning + Acting) pattern for AI agents: how to combine chain-of-thought reasoning with tool use for better agent performance.

Last updated

Core9 min readFirst readThe Agent Loop

After this section you can

  • Trace a ReAct episode through thought, action and observation, and say what each step depends on
  • State what the ReAct paper showed, including where it did not win on its own
  • Map text-format ReAct onto native tool use and thinking, and diagnose ReAct-specific failures
10

ReAct: Reasoning + Acting in One Loop

The 2022 paper that had models think between tool calls, and the pattern that tool-use APIs later built in. What it showed, what a trace looks like, and how it fails.

Key idea

ReAct interleaves a written thought before each action and reads the observation before the next thought. Reasoning decides what to look up; what comes back corrects the reasoning. Each step depends on evidence the model did not have one step earlier.

Reasoning alone guesses, acting alone flails. ReAct alternates them in one transcript.
REASON ONLY (CoT) Question Thought → thought → thought from what the model memorised Answer no new facts, so it can hallucinate ACT ONLY Question Action → observation → action no reasoning between calls Answer no plan, so it cannot decompose or recover REACT Question Thought → Action → Observation each thought reads the last result Answer grounded, and the plan can change mid-task repeat

One episode, line by line

ReAct started as a prompting format: labelled lines in one growing transcript. The model writes a Thought and an Action; the runtime runs the action and appends the Observation.

A multi-hop question, ReAct style
user

Aside from the Apple Remote, what other device can control the program the Apple Remote was originally designed to interact with?

assistant

Thought 1: I need the program the Apple Remote was designed for.
Action 1: search[Apple Remote]

tool

Observation 1: The Apple Remote … was designed to control the Front Row media center program …

The fact that decides the next step arrives here.
assistant

Thought 2: The program is Front Row. What else controls Front Row?
Action 2: search[Front Row (software)]

This thought could not have been written before Observation 1.
tool

Observation 2: Front Row … is controlled by an Apple Remote or the keyboard function keys.

assistant

Thought 3: The other device is the keyboard function keys.
Action 3: finish[keyboard function keys]

Adapted from the paper’s opening HotpotQA example; observations abridged.

Why interleaving works

The paper’s claim is a two-way benefit. Neither half does the job alone, which is what its reason-only and act-only baselines were built to show.

Reasoning helps acting thought → action
  • breaks the goal into lookups
  • tracks what is known and what is still missing
  • notices a surprise and changes plan
Acting helps reasoning observation → thought
  • brings in facts the model never memorised
  • replaces a guess with a lookup
  • leaves a readable trail of what was checked

Interactive: step through a ReAct episode that looks up an account, reads its spend and files a ticket.

What the paper showed

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (posted October 2022, ICLR 2023). The main results prompt PaLM-540B with a handful of worked examples and no fine-tuning, and give it a simple Wikipedia API (search, lookup) for the knowledge tasks.

TaskWhat it testsResult
HotpotQAmulti-hop questions, exact matchReAct 27.4 vs chain-of-thought 29.4; ReAct falling back to self-consistent CoT: 35.1
FEVERfact verification, accuracyReAct 60.9 vs chain-of-thought 56.3; the best combination: 64.6
ALFWorldtext-based household tasks+34 points absolute success rate over imitation and RL baselines
WebShopshopping on a simulated site+10 points absolute success rate over those baselines

The honest reading: on HotpotQA, ReAct alone did not beat chain-of-thought. It failed differently: chain-of-thought mostly by inventing facts, ReAct through unhelpful searches and repeated steps that show up in the trace. Combining the two scored best.

On ALFWorld and WebShop the prompt held only one or two examples, yet it beat methods trained on the task.

The text loop and the stop-sequence trick

Before native tool calling, the runtime parsed actions out of text. One line did the safety work: stop generation at the word Observation, or the model writes its own observations and reasons over fiction.

SYSTEM = """Interleave Thought, Action and Observation lines. Write one Thought
and one Action, then stop. Actions: search[entity], lookup[keyword], finish[answer]."""
ACTION = re.compile(r"Action \d+: (\w+)\[(.*?)\]")

def react_text(question, llm, tools, max_steps=8):
    transcript = f"Question: {question}\n"
    for n in range(1, max_steps + 1):
        out = llm(SYSTEM, transcript, stop=[f"\nObservation {n}:"])  # the trick
        transcript += out
        m = ACTION.search(out)
        if m and m.group(1) == "finish":
            return m.group(2)
        tool = tools.get(m.group(1)) if m else None
        obs = tool(m.group(2))[:1500] if tool else "Invalid action. Use search, lookup or finish."
        transcript += f"\nObservation {n}: {obs}\n"
    return None                       # step cap hit: report it, do not guess

Still the pattern for models without native tool calling. The loop mechanics and guards are the same as in The Agent Loop.

How ReAct became native

Tool-use APIs absorbed the pattern. You no longer prompt for labelled lines or parse them; each piece has a structured home.

ReAct in 2022 (text)Native tool use today
Thought: free text in the promptthinking between tool calls; on current Claude models, adaptive thinking (see Adaptive Thinking & Effort)
Action: search[query] matched by a regexa tool_use block with JSON input checked against a schema
Stop sequence at “Observation”stop_reason: "tool_use"; the model stops itself
Observation: appended texta tool_result block keyed by tool_use_id
The transcriptthe messages array

One difference matters for audits. On current Claude models the raw reasoning is never returned: thinking text is omitted by default, and display: "summarized" gives a summary. Pass thinking blocks back unchanged, and treat the tool_use and tool_result pairs as the exact record of what the agent did.

Failure modes specific to ReAct

FailureWhat the trace showsFix
Invented observationan “Observation” the runtime never wrote (text format)stop sequence; native tool use rules it out
Repetitionthe same thought and search, lightly rewordeddetect repeats, say what was already tried, cap steps
Derailed by a bad resultan unhelpful search sends every later thought off coursea search tool that returns titles and snippets; allow a fallback to answering from knowledge, as the paper’s combination did
Malformed actionthe parser misses, plan text is taken as the answerstrict parsing plus a corrective observation, or native tools
Overthinking easy asksfour searches for a fact the model knowsroute simple questions to one call; lower effort
Interview angle

“What is ReAct and do you still use it?” Say why it works: each thought is conditioned on fresh evidence, so the plan can change mid-task. Then say where it lives now: native tool use plus thinking is ReAct built into the API, so you write explicit Thought/Action prompts only for models without tool calling. Mention the stop-sequence trick to show you have run one.

Check yourself
A text-format ReAct agent logs “Observation: payment succeeded”, but the payment API shows no call. What happened, and what is the fix?

The model wrote the observation itself because generation did not stop after the Action. Stop at the Observation label so only the runtime writes observations, or move to native tool use, where results can only come from your code.

On HotpotQA, ReAct alone scored below chain-of-thought on exact match. Why is that not a case against ReAct?

The two failed differently: chain-of-thought mostly by inventing facts, ReAct through unhelpful searches and repeated steps. Combining them scored best, and ReAct’s errors were visible in the trace instead of hidden in fluent prose.

You move a ReAct prompt onto Claude with native tools and adaptive thinking. What do you delete from the old prompt, and what do you keep in your logs?

Delete the Thought/Action/Observation format, the action grammar and the regex parsing: tool definitions, tool_use blocks and thinking replace them. Log the tool_use and tool_result pairs as the exact record; thinking comes back as a summary at most.

Recap
  • ReAct alternates Thought, Action and Observation so reasoning stays grounded and plans can change.
  • The paper: grounded answers on HotpotQA and FEVER, best when combined with chain-of-thought; +34 and +10 points on ALFWorld and WebShop from a prompt with one or two examples.
  • In text form, a stop sequence at “Observation” keeps the model from inventing results.
  • Native tool use plus thinking is ReAct built into the API; the tool calls and results are the exact record.

Next: Reflexion: Agents That Learn from Mistakes, what an agent does after a ReAct episode fails.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium