ReAct Pattern
ReAct (Reasoning + Acting) pattern for AI agents: how to combine chain-of-thought reasoning with tool use for better agent performance.
Last updated
After this section you can
- Trace a ReAct episode through thought, action and observation, and say what each step depends on
- State what the ReAct paper showed, including where it did not win on its own
- Map text-format ReAct onto native tool use and thinking, and diagnose ReAct-specific failures
ReAct: Reasoning + Acting in One Loop
The 2022 paper that had models think between tool calls, and the pattern that tool-use APIs later built in. What it showed, what a trace looks like, and how it fails.
ReAct interleaves a written thought before each action and reads the observation before the next thought. Reasoning decides what to look up; what comes back corrects the reasoning. Each step depends on evidence the model did not have one step earlier.
One episode, line by line
ReAct started as a prompting format: labelled lines in one growing transcript. The model writes a Thought and an Action; the runtime runs the action and appends the Observation.
Aside from the Apple Remote, what other device can control the program the Apple Remote was originally designed to interact with?
Thought 1: I need the program the Apple Remote was designed for.
Action 1: search[Apple Remote]
Observation 1: The Apple Remote … was designed to control the Front Row media center program …
The fact that decides the next step arrives here.Thought 2: The program is Front Row. What else controls Front Row?
Action 2: search[Front Row (software)]
Observation 2: Front Row … is controlled by an Apple Remote or the keyboard function keys.
Thought 3: The other device is the keyboard function keys.
Action 3: finish[keyboard function keys]
Adapted from the paper’s opening HotpotQA example; observations abridged.
Why interleaving works
The paper’s claim is a two-way benefit. Neither half does the job alone, which is what its reason-only and act-only baselines were built to show.
- breaks the goal into lookups
- tracks what is known and what is still missing
- notices a surprise and changes plan
- brings in facts the model never memorised
- replaces a guess with a lookup
- leaves a readable trail of what was checked
Interactive: step through a ReAct episode that looks up an account, reads its spend and files a ticket.
What the paper showed
Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (posted October 2022, ICLR 2023). The main results prompt PaLM-540B with a handful of worked examples and no fine-tuning, and give it a simple Wikipedia API (search, lookup) for the knowledge tasks.
| Task | What it tests | Result |
|---|---|---|
| HotpotQA | multi-hop questions, exact match | ReAct 27.4 vs chain-of-thought 29.4; ReAct falling back to self-consistent CoT: 35.1 |
| FEVER | fact verification, accuracy | ReAct 60.9 vs chain-of-thought 56.3; the best combination: 64.6 |
| ALFWorld | text-based household tasks | +34 points absolute success rate over imitation and RL baselines |
| WebShop | shopping on a simulated site | +10 points absolute success rate over those baselines |
The honest reading: on HotpotQA, ReAct alone did not beat chain-of-thought. It failed differently: chain-of-thought mostly by inventing facts, ReAct through unhelpful searches and repeated steps that show up in the trace. Combining the two scored best.
On ALFWorld and WebShop the prompt held only one or two examples, yet it beat methods trained on the task.
The text loop and the stop-sequence trick
Before native tool calling, the runtime parsed actions out of text. One line did the safety work: stop generation at the word Observation, or the model writes its own observations and reasons over fiction.
SYSTEM = """Interleave Thought, Action and Observation lines. Write one Thought
and one Action, then stop. Actions: search[entity], lookup[keyword], finish[answer]."""
ACTION = re.compile(r"Action \d+: (\w+)\[(.*?)\]")
def react_text(question, llm, tools, max_steps=8):
transcript = f"Question: {question}\n"
for n in range(1, max_steps + 1):
out = llm(SYSTEM, transcript, stop=[f"\nObservation {n}:"]) # the trick
transcript += out
m = ACTION.search(out)
if m and m.group(1) == "finish":
return m.group(2)
tool = tools.get(m.group(1)) if m else None
obs = tool(m.group(2))[:1500] if tool else "Invalid action. Use search, lookup or finish."
transcript += f"\nObservation {n}: {obs}\n"
return None # step cap hit: report it, do not guess
Still the pattern for models without native tool calling. The loop mechanics and guards are the same as in The Agent Loop.
How ReAct became native
Tool-use APIs absorbed the pattern. You no longer prompt for labelled lines or parse them; each piece has a structured home.
| ReAct in 2022 (text) | Native tool use today |
|---|---|
| Thought: free text in the prompt | thinking between tool calls; on current Claude models, adaptive thinking (see Adaptive Thinking & Effort) |
Action: search[query] matched by a regex | a tool_use block with JSON input checked against a schema |
| Stop sequence at “Observation” | stop_reason: "tool_use"; the model stops itself |
| Observation: appended text | a tool_result block keyed by tool_use_id |
| The transcript | the messages array |
One difference matters for audits. On current Claude models the raw reasoning is never returned: thinking text is omitted by default, and display: "summarized" gives a summary. Pass thinking blocks back unchanged, and treat the tool_use and tool_result pairs as the exact record of what the agent did.
Failure modes specific to ReAct
| Failure | What the trace shows | Fix |
|---|---|---|
| Invented observation | an “Observation” the runtime never wrote (text format) | stop sequence; native tool use rules it out |
| Repetition | the same thought and search, lightly reworded | detect repeats, say what was already tried, cap steps |
| Derailed by a bad result | an unhelpful search sends every later thought off course | a search tool that returns titles and snippets; allow a fallback to answering from knowledge, as the paper’s combination did |
| Malformed action | the parser misses, plan text is taken as the answer | strict parsing plus a corrective observation, or native tools |
| Overthinking easy asks | four searches for a fact the model knows | route simple questions to one call; lower effort |
“What is ReAct and do you still use it?” Say why it works: each thought is conditioned on fresh evidence, so the plan can change mid-task. Then say where it lives now: native tool use plus thinking is ReAct built into the API, so you write explicit Thought/Action prompts only for models without tool calling. Mention the stop-sequence trick to show you have run one.
A text-format ReAct agent logs “Observation: payment succeeded”, but the payment API shows no call. What happened, and what is the fix?
The model wrote the observation itself because generation did not stop after the Action. Stop at the Observation label so only the runtime writes observations, or move to native tool use, where results can only come from your code.
On HotpotQA, ReAct alone scored below chain-of-thought on exact match. Why is that not a case against ReAct?
The two failed differently: chain-of-thought mostly by inventing facts, ReAct through unhelpful searches and repeated steps. Combining them scored best, and ReAct’s errors were visible in the trace instead of hidden in fluent prose.
You move a ReAct prompt onto Claude with native tools and adaptive thinking. What do you delete from the old prompt, and what do you keep in your logs?
Delete the Thought/Action/Observation format, the action grammar and the regex parsing: tool definitions, tool_use blocks and thinking replace them. Log the tool_use and tool_result pairs as the exact record; thinking comes back as a summary at most.
- ReAct alternates Thought, Action and Observation so reasoning stays grounded and plans can change.
- The paper: grounded answers on HotpotQA and FEVER, best when combined with chain-of-thought; +34 and +10 points on ALFWorld and WebShop from a prompt with one or two examples.
- In text form, a stop sequence at “Observation” keeps the model from inventing results.
- Native tool use plus thinking is ReAct built into the API; the tool calls and results are the exact record.