English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ReasoningBank: Teaching AI Agents to Learn from Mistakes with Reasoning Memory

Forum topic · 小凯 · 2026-05-28

Summary

ReasoningBank (ICLR 2026, Google Research) is a memory framework that lets LLM-based agents stop repeating the same mistakes. Instead of storing raw trajectories or successful workflows like Synapse or AWM, it stores distilled reasoning strategies in three fields: title, description, and content. Crucially, it also stores lessons from failures as counterfactual signals, improving WebArena success rates by +3.2 points where baselines gain nothing or regress. The companion method MaTTS (Memory-aware Test-Time Scaling) shows memory is a prerequisite for test-time scaling to work: with memory, scaling from k=1 to k=5 yields +16.1 points versus +1.6 without. Experiments across WebArena, Mind2Web, and SWE-Bench-Verified show consistent gains (e.g., Gemini-2.5-pro reaching 56.3 on WebArena with MaTTS), 26.9% step efficiency improvement, robustness to imperfect LLM-as-a-Judge labels, and emergent evolution of strategies from low-level actions to adaptive, composable policies. Code is available on GitHub.

Your Agent Repeats the Same Mistake Every Time — ReasoningBank Teaches It to Remember Lessons

> Source: ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory, ICLR 2026, https://arxiv.org/pdf/2509.25140

---

1. Introduction: Agent Amnesia

Ask Claude Code to fix a bug. It spends 20 minutes, tries three approaches, and finds the right one. Three days later, a similar bug appears. It starts from scratch and walks through all three approaches again.

The agent isn't dumb — it just has no memory.

Existing memory approaches either store raw trajectories (verbose, noisy) or successful workflows (ignoring lessons hidden in failures). ReasoningBank's insight: what should be stored is not "what was done" but "why it was done" — the reasoning strategy itself.

---

2. Core Design: Store Strategies, Not Trajectories

The memory structure is minimal, with just three fields:

| Field | Purpose | Example | |-------|---------|---------| | Title | Strategy identifier | "Prioritize user account sections for personal data" | | Description | One-sentence summary | "When a query requests user-specific information..." | | Content | Distilled reasoning steps | "Systematically look for and click on links..." |

The key is the abstraction level: not "clicked this button," but "when handling user-data queries, prioritize account-related sections." Such strategies transfer across tasks and domains.

---

3. The Value of Failure: Counterfactual Signals

The paper's most counterintuitive finding: failed trajectories are more instructive than successful ones.

Prior methods (Synapse, AWM) only store successes. ReasoningBank stores both:

  • Successes → validated effective strategies
  • Failures → counterfactual signals and pitfall avoidance
  • Hard experimental data (Figure 7):

    | Method | Success-only | With failures | Change | |--------|-------------|---------------|--------| | Synapse | 40.6 | 41.7 | +1.1 (negligible) | | AWM | 44.4 | 42.2 | -2.2 (drops!) | | ReasoningBank | 46.5 | 49.7 | +3.2 (significant gain) |

    Synapse and AWM cannot exploit failures because their formats are designed for "success procedures." ReasoningBank's abstract strategy format naturally accommodates failure — "don't do this next time" is also a strategy.

    ---

    4. MaTTS: Memory-aware Test-Time Scaling

    The paper also combines memory with Test-Time Scaling (TTS).

    Traditional TTS works poorly for agents because each sample explores from scratch with no accumulation — scaling without memory just means "repeating mistakes more times."

    MaTTS creates a flywheel: 1. Use memory strategies for better initial sampling 2. Multiple samples produce more successes/failures 3. Successes enrich the memory library 4. Better memory → better sampling next round

    Scaling curves (Figure 4, WebArena-Shopping):

    | k | Without memory | With memory | |---|----------------|-------------| | 1 | 39.0 | 49.7 | | 3 | 42.2 | 52.9 | | 5 | 40.6 | 55.1 |

    Without memory, k=5 gains only 1.6 over k=1. With memory, +16.1. Memory is the prerequisite for scaling to work.

    ---

    5. Emergence: Strategies Evolve from Low-Level to High-Level

    The paper traces memory items' evolution (Figure 6):

    | Stage | Strategy type | Example | |-------|---------------|---------| | Early | Execution-oriented | "find navigation links", "click on 'Next Page'" | | Mid | Self-reflection | "re-verifying identifiers to reduce simple mistakes" | | Later | Adaptive checks | "systematically leverage search or filters to ensure completeness" | | Mature | Compositional | "cross-referencing task requirements and reassessing options" |

    The agent wasn't taught these strategies — it distilled them itself while solving problems. ReasoningBank provides the container for storing and retrieving them, letting strategies accumulate, get reused, and evolve.

    ---

    6. Experiments: Gains Across Three Benchmarks

    WebArena (web browsing, 684 tasks):

    | Model | No memory | ReasoningBank | Gain | |-------|-----------|---------------|------| | Gemini-2.5-flash | 40.5 | 48.8 | +8.3 | | Gemini-2.5-pro | 46.7 | 53.9 | +7.2 | | Claude-3.7-sonnet | 41.7 | 46.3 | +4.6 |

    With MaTTS: Gemini-2.5-flash reaches 51.8, Gemini-2.5-pro reaches 56.3.

    Step efficiency: Shopping tasks drop from 8.2 to 6.1 steps (-26.9%).

    Mind2Web (cross-domain generalization):

  • Cross-Task: 3.3 → 4.8
  • Cross-Website: 3.4 → 3.8
  • Cross-Domain: 1.0 → 1.6
  • SWE-Bench-Verified (software engineering):

  • Gemini-2.5-flash: 34.2 → 38.8
  • Gemini-2.5-pro: 54.0 → 57.4
  • ---

    7. Robustness to Judge Accuracy

    The framework uses LLM-as-a-Judge to label success/failure, which introduces errors:

    | Judge accuracy | Success rate | |----------------|--------------| | 100% (ground truth) | 49.7 | | 90% | ~49.4 | | 70% | ~48.2 | | 50% (random) | ~47.6 |

    Within the realistic 70–90% error range, performance barely changes. ReasoningBank is robust to verification noise — no perfect labels required.

    ---

    8. Comparison with Existing Memory Methods

    | Dimension | Synapse | AWM | ReasoningBank | |-----------|---------|-----|---------------| | What is stored | Raw trajectories | Success procedures | Reasoning strategies | | Abstraction level | Low | Medium | High | | Uses failures | No | No | Yes | | Cross-task transfer | Weak | Medium | Strong | | Human-readable | Poor | Medium | Good | | Direct prompt injection | Needs parsing | Needs matching | Directly usable |

    ---

    9. Conclusion: Agent Self-Evolution

    ReasoningBank's core insight: the scaling bottleneck for agents is not compute — it's experience. Not "doing more tasks" (breadth scaling) but "doing each task thoroughly" (depth scaling).

    Every failure is not waste — it's data. Every success is not an endpoint — it's a strategy. When an agent can store these strategies, retrieve them, and inject them into new tasks, it begins to self-evolve.

    This isn't future tense. Gemini-2.5-flash + ReasoningBank already reaches 56.3% on WebArena — half a year ago, SOTA on this benchmark was below 40%.

    > "An agent shouldn't start from zero every time. It should remember how it learned."

    ---

    References

  • ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory, ICLR 2026, https://arxiv.org/pdf/2509.25140
  • Code: https://github.com/google-research/reasoning-bank

Tags

#reasoningbank#ai-agents#agent-memory#test-time-scaling#matts#webarena#self-evolving-agents#google-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980472