What problem does ReAct solve?
Traditional Chain-of-Thought prompting asks a large language model to produce a full reasoning trace first and then commit to a final answer. If the model realizes mid-thought that it is missing a fact, it cannot revise its plan because the chain has already been written. ReAct (Yao et al., ICLR 2023, https://openreview.net/forum?id=WE_vluYUL-X) proposes instead to interleave reasoning and acting: every Thought is followed by an Action, every Action yields an Observation, and the next Thought conditions on that Observation.
How the ReAct loop works
The model's output alternates between three modes:
1. Thought — "I need to look up the capital of France."
2. Action — call a search tool with query "capital of France".
3. Observation — the tool returns "Paris".
4. Thought — "The capital is Paris, so I can answer now."
5. Action — emit the final answer.
This loop repeats until the model produces a terminating action. Training is pure few-shot prompting: a handful of ReAct trajectories in the prompt teach the alternating pattern, no fine-tuning required.
Why interleaving beats separation
On a toy multi-step math problem the two approaches look almost identical. The difference becomes decisive when external knowledge is required. Asked "What is the main contribution of the 2024 Nobel Physics laureate?", a pure CoT model may hallucinate. ReAct first emits a Thought about needing evidence, issues a search Action, receives an Observation, and only then continues reasoning. Reasoning is no longer an isolated mental act — it is a continuous dialogue with the outside world.
Key empirical findings
- Knowledge-intensive QA. ReAct significantly outperforms CoT when answers depend on facts that may be absent from the model's parametric memory.
- Multi-step decision making. On interactive environments such as ALFWorld, ReAct's grounded trajectories yield higher success rates.
- Interpretability. The Thought-Action-Observation trace is explicit and human-readable, exposing the model's intermediate decisions.
- Error propagation. Thought quality directly determines Action quality: a wrong thought triggers a wrong action. The paper mitigates this by forcing the model to emit a Thought at every step, accepting some overhead in exchange for reliability.
- Toolformer — learned API calling.
- AutoGPT — autonomous multi-step task execution.
- LangChain — standardized LLM + tool orchestration.
- Deep Research systems — iterative search-and-synthesis pipelines.
- Predefined action space. Only a fixed set of actions (Search, Lookup, Finish, etc.) is available. Problems requiring novel tools fall outside the loop. Later work such as Toolformer enlarges the action space but raises the difficulty of picking the right action.
- No termination guarantee. The model can loop indefinitely between Thought and Action. The paper imposes a hard step limit — a brute-force constraint rather than an elegant solution.
- Failure handling. When an Action fails or returns garbage, ReAct has limited mechanisms for recovery beyond hoping the next Thought is correct.
The bigger insight
ReAct's lasting contribution is conceptual rather than mechanical. Before ReAct, LLMs were closed systems that mapped text-in to text-out. After ReAct, they became open systems that can proactively invoke tools, query databases, browse the web, and adapt their behavior to external feedback. This paradigm shift underlies essentially every subsequent Agentic LLM framework:
ReAct opened a window between the LLM and the world; earlier models could only look through the glass, later models can reach through it.
Critical perspective
ReAct has structural limits the original paper does not deeply address:
Takeaways
ReAct establishes three enduring principles for Agentic LLMs:
1. Reasoning and acting should be interleaved, not separated. 2. LLMs can initiate interaction with the external world. 3. Feedback from those interactions must flow back into reasoning.
It also hands down the open problems — scaling action spaces, ensuring termination, and recovering from tool failures — that the field continues to wrestle with. For anyone tracing the lineage of Agentic AI, ReAct is the unavoidable starting point.
> "You may know the name of that bird in every language, yet you still know nothing about the bird." ReAct does more than name the bird — it lets us reach out and feel its feathers.