1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Reflexion

Reflexion pattern for self-improving AI agents: verbal reinforcement learning where agents turn failures into written lessons and retry — no gradients required.

Last updated

Core10 min readFirst readThe Agent LoopReAct Pattern

After this section you can

  • Implement the evaluate, reflect and retry loop with a capped lesson memory
  • Choose an evaluator for a task and explain why a weak one makes retries harmful
  • Price a retry budget per task and tell Reflexion apart from self-refine
11

Reflexion: Agents That Learn from Mistakes

The agent fails, a real check says why, it writes itself a lesson and retries with that lesson in its prompt. Learning in plain text within a task, with no fine-tuning, and only as good as the check that drives it.

Key idea

Reflexion turns a failure signal into a written lesson and carries it into the next attempt. The lesson is only as good as the signal: with a trustworthy evaluator such as tests, retries rescue real failures; with a weak one, the agent confidently fixes the wrong thing.

Fail, get a real signal, write the lesson down, try again with it. No weights change.
1 · Actor attempts the task: trial n often a whole ReAct episode 2 · Evaluator tests · exact match · heuristics the load-bearing part pass return the output and the trial count 3 · Self-reflection why did it fail? what changes next time? 4 · Lesson memory the last few lessons plain text, capped retry cap reached stop and report the failure, with the lessons, to a human output passes fails + feedback (the error text) lesson added to trial n + 1 Between trials the weights stay fixed; only the prompt changes, and it carries the lessons.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium