Your Agent Repeats the Same Mistake Every Time — ReasoningBank Teaches It to Remember Lessons
> Source: ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory, ICLR 2026, https://arxiv.org/pdf/2509.25140
---
1. Introduction: Agent Amnesia
Ask Claude Code to fix a bug. It spends 20 minutes, tries three approaches, and finds the right one. Three days later, a similar bug appears. It starts from scratch and walks through all three approaches again.
The agent isn't dumb — it just has no memory.
Existing memory approaches either store raw trajectories (verbose, noisy) or successful workflows (ignoring lessons hidden in failures). ReasoningBank's insight: what should be stored is not "what was done" but "why it was done" — the reasoning strategy itself.
---
2. Core Design: Store Strategies, Not Trajectories
The memory structure is minimal, with just three fields:
| Field | Purpose | Example | |-------|---------|---------| | Title | Strategy identifier | "Prioritize user account sections for personal data" | | Description | One-sentence summary | "When a query requests user-specific information..." | | Content | Distilled reasoning steps | "Systematically look for and click on links..." |
The key is the abstraction level: not "clicked this button," but "when handling user-data queries, prioritize account-related sections." Such strategies transfer across tasks and domains.
---
3. The Value of Failure: Counterfactual Signals
The paper's most counterintuitive finding: failed trajectories are more instructive than successful ones.
Prior methods (Synapse, AWM) only store successes. ReasoningBank stores both:
- Successes → validated effective strategies
- Failures → counterfactual signals and pitfall avoidance
- Cross-Task: 3.3 → 4.8
- Cross-Website: 3.4 → 3.8
- Cross-Domain: 1.0 → 1.6
- Gemini-2.5-flash: 34.2 → 38.8
- Gemini-2.5-pro: 54.0 → 57.4
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory, ICLR 2026, https://arxiv.org/pdf/2509.25140
- Code: https://github.com/google-research/reasoning-bank
Hard experimental data (Figure 7):
| Method | Success-only | With failures | Change | |--------|-------------|---------------|--------| | Synapse | 40.6 | 41.7 | +1.1 (negligible) | | AWM | 44.4 | 42.2 | -2.2 (drops!) | | ReasoningBank | 46.5 | 49.7 | +3.2 (significant gain) |
Synapse and AWM cannot exploit failures because their formats are designed for "success procedures." ReasoningBank's abstract strategy format naturally accommodates failure — "don't do this next time" is also a strategy.
---
4. MaTTS: Memory-aware Test-Time Scaling
The paper also combines memory with Test-Time Scaling (TTS).
Traditional TTS works poorly for agents because each sample explores from scratch with no accumulation — scaling without memory just means "repeating mistakes more times."
MaTTS creates a flywheel: 1. Use memory strategies for better initial sampling 2. Multiple samples produce more successes/failures 3. Successes enrich the memory library 4. Better memory → better sampling next round
Scaling curves (Figure 4, WebArena-Shopping):
| k | Without memory | With memory | |---|----------------|-------------| | 1 | 39.0 | 49.7 | | 3 | 42.2 | 52.9 | | 5 | 40.6 | 55.1 |
Without memory, k=5 gains only 1.6 over k=1. With memory, +16.1. Memory is the prerequisite for scaling to work.
---
5. Emergence: Strategies Evolve from Low-Level to High-Level
The paper traces memory items' evolution (Figure 6):
| Stage | Strategy type | Example | |-------|---------------|---------| | Early | Execution-oriented | "find navigation links", "click on 'Next Page'" | | Mid | Self-reflection | "re-verifying identifiers to reduce simple mistakes" | | Later | Adaptive checks | "systematically leverage search or filters to ensure completeness" | | Mature | Compositional | "cross-referencing task requirements and reassessing options" |
The agent wasn't taught these strategies — it distilled them itself while solving problems. ReasoningBank provides the container for storing and retrieving them, letting strategies accumulate, get reused, and evolve.
---
6. Experiments: Gains Across Three Benchmarks
WebArena (web browsing, 684 tasks):
| Model | No memory | ReasoningBank | Gain | |-------|-----------|---------------|------| | Gemini-2.5-flash | 40.5 | 48.8 | +8.3 | | Gemini-2.5-pro | 46.7 | 53.9 | +7.2 | | Claude-3.7-sonnet | 41.7 | 46.3 | +4.6 |
With MaTTS: Gemini-2.5-flash reaches 51.8, Gemini-2.5-pro reaches 56.3.
Step efficiency: Shopping tasks drop from 8.2 to 6.1 steps (-26.9%).
Mind2Web (cross-domain generalization):
SWE-Bench-Verified (software engineering):
---
7. Robustness to Judge Accuracy
The framework uses LLM-as-a-Judge to label success/failure, which introduces errors:
| Judge accuracy | Success rate | |----------------|--------------| | 100% (ground truth) | 49.7 | | 90% | ~49.4 | | 70% | ~48.2 | | 50% (random) | ~47.6 |
Within the realistic 70–90% error range, performance barely changes. ReasoningBank is robust to verification noise — no perfect labels required.
---
8. Comparison with Existing Memory Methods
| Dimension | Synapse | AWM | ReasoningBank | |-----------|---------|-----|---------------| | What is stored | Raw trajectories | Success procedures | Reasoning strategies | | Abstraction level | Low | Medium | High | | Uses failures | No | No | Yes | | Cross-task transfer | Weak | Medium | Strong | | Human-readable | Poor | Medium | Good | | Direct prompt injection | Needs parsing | Needs matching | Directly usable |
---
9. Conclusion: Agent Self-Evolution
ReasoningBank's core insight: the scaling bottleneck for agents is not compute — it's experience. Not "doing more tasks" (breadth scaling) but "doing each task thoroughly" (depth scaling).
Every failure is not waste — it's data. Every success is not an endpoint — it's a strategy. When an agent can store these strategies, retrieve them, and inject them into new tasks, it begins to self-evolve.
This isn't future tense. Gemini-2.5-flash + ReasoningBank already reaches 56.3% on WebArena — half a year ago, SOTA on this benchmark was below 40%.
> "An agent shouldn't start from zero every time. It should remember how it learned."
---
References