Have you ever fixed a bug after two hours, then spent another two hours on a similar bug later? The problem isn't forgetting how you fixed it — it's failing to recognize that *this* bug and *that* bug are the same kind of problem. Human memory retrieves by causal structure, not timeline: "last time it was a thread-safety issue, this one looks like that too."
AI Agents have the same flaw. Most current agent repair systems are stateless: when an error appears, they start from scratch with no memory of similar past errors or dead ends. Adding memory? Prior approaches either store the last N conversations (too crude) or use semantic retrieval (losing causal structure).
The August 2026 paper *Causal Episodic Memory for Feedback-Driven Agent Repair* proposes MERIT (Memory-Augmented Error-Typed Retrieval) — an "experience hard drive" for agents: stored by error type, retrieved by causal structure.
Paper: https://arxiv.org/abs/2608.05906
1. The Problem: Stateless vs. Memory-Augmented Repair
Text-to-SQL: An Ideal Test Bed
The paper tests on Text-to-SQL: the user speaks in natural language ("show stores with sales over 100K last month"), and the Agent translates it into SQL. This scenario works well because:
1. Errors are classifiable: syntax errors, wrong table names, wrong columns, JOIN errors, aggregation errors, etc. 2. Repairs are verifiable: run the SQL and you know if it's correct. 3. Errors repeat: the same error type recurs across different queries.
Limits of Stateless Repair
Existing Text-to-SQL repair methods are mostly stateless — regenerate on error, or use dynamic RAG against documentation. Problems:
- No memory of past fixes: every "GROUP BY missing aggregate column" error is reasoned about from scratch.
- No memory of dead ends: the agent may retry the same plausible-but-failing direction.
- No error-type awareness: one strategy for all error types.
- Positive Memory: Oracle-verified correct fixes — "for error type X, fix Y worked."
- Negative Memory: attempted directions that failed — "for error type X, fix Z didn't work."
Memory Exists, But Not Good Enough
Prior memory methods like Dynamic RAG store past repairs and semantically retrieve similar cases, but:
1. Coarse retrieval granularity: semantic similarity can surface "looks similar but actually a different error type" cases. 2. No distinction of positive vs. negative experience: only successes are stored. But "this path doesn't work" is also valuable.
MERIT's切入点 is exactly these two points: store by error type + keep both positive and negative experience.
2. MERIT's Core Design
Typed Positive and Negative Memory
MERIT's memory is structured, not a flat case list:
Type-Conditioned Retrieval
Retrieval first filters by error type, then does semantic search within the type. This avoids cross-type mistakes: a syntax-error fix won't be retrieved for a logic error even if their descriptions read similarly.
Causal Temporal Constraint
"Causal" has a specific meaning here: only completed episodes enter memory. An in-progress repair cannot be a memory source — you don't yet know if it worked. Just like humans: you don't present an unverified attempt as experience; only validated lessons earn an "I've seen this before."
3. Results: Honest Good News and Bad News
| Method | Spider (EX) | BIRD (EX) | |--------|-------------|-----------| | Iterative repair baseline | 66.34% | 47.35% | | MERIT | 69.79% | 48.44% | | Dynamic RAG | 69.51% | 48.16% |
MERIT gains +3.45pp on Spider and +1.09pp on BIRD over the stateless baseline.
The Honest Bad News: It Didn't Beat Dynamic RAG
Look closer: the gap vs. Dynamic RAG is 0.28pp on both benchmarks — within statistical noise; MERIT cannot be said to reliably outperform Dynamic RAG. The paper admits this, which is commendable — many papers would cherry-pick angles to claim superiority.
Why Didn't MERIT Significantly Beat Dynamic RAG?
1. Error classifier accuracy is only 48%: nearly half the time, type-conditioned retrieval uses the wrong type, actively misleading retrieval. 2. Query order sensitivity: memory grows dynamically; earlier queries shape later retrievals, so results vary with order. 3. Not cheaper than stateless methods: maintaining and querying memory has overhead without clearly proportional benefit.
4. Why It Still Matters
1. Conceptual contribution over numeric contribution: MERIT opens a design space — error type × positive/negative experience × temporal constraint — even if the current implementation doesn't win. 2. Negative experience is an original contribution: previous memory-augmented methods stored only successes. MERIT is the first to systematically store and use failure directions — closer to how human experts actually work: knowing "what not to do." 3. Honest reporting is itself a contribution: in a benchmark-chasing field, admitting "no significant improvement over baseline" prevents later researchers from stepping in the same potholes.
5. Conceptual Positioning: Granularity Isomorphism
MERIT instantiates a "granularity isomorphism" principle: the optimization granularity should match the granularity of the thing being optimized. Prior examples: Heddle and CodeRescue raised decision granularity from single calls to trajectories/recovery actions; ACE's RPI workflow cut context management from whole-conversation to per-phase; Regression Tax refined evaluation from average pass rate to paired structure. MERIT refines memory granularity from case-level to error type × polarity — isomorphic to Heddle's trajectory-level decisions.
6. Relation to "Changing the Level of Problem-Solving"
MERIT also belongs to the "solve at a different level" lineage: octopus RNA editing (change RNA, not DNA), slime mold externalized memory (mucus trails, not neurons), MACRO (layer execution paths, not weights). MERIT's level switch: don't fine-tune or retrain the model — externalize experience into a memory store. But its honest results also remind us level-switching isn't always effective; with a weak classifier, the benefit can drown in noise.
7. Limitations and Future Directions
1. The error classifier is the bottleneck: at 48% accuracy, type-conditioned retrieval is wrong almost half the time. Raising it to 80%+ could let MERIT clearly beat Dynamic RAG. 2. Only tested on Text-to-SQL: error types here are well-structured; transfer to code generation or math reasoning is unverified. 3. Cold start: an empty memory offers no help early on; no cold-start strategy is discussed. 4. Capacity management: unbounded memory growth raises retrieval cost; no forgetting or compression strategy is discussed.
The most promising future direction: improving error classifier accuracy. Like giving a veteran programmer a good diagnostic tool — only then does the "experience hard drive" truly pay off.
8. Conclusion
MERIT gives agents an "experience hard drive" — stored by error type, retrieved by causal structure. It doesn't significantly beat Dynamic RAG yet, but the typed positive/negative memory framework points the way forward. The most memorable takeaway isn't the numbers but a design principle: memory isn't just about what to store, but how to store it and when it can be retrieved. Error type × outcome polarity × temporal constraint define a structured design space.
As an old engineer's saying goes: "Novices remember answers; experts remember potholes." MERIT tries to move agents from "only remembering answers" to "also remembering potholes." The gains aren't there yet, but the direction is right — after all, knowing "what not to do" is often worth more than knowing "what to do."
---
Paper link: https://arxiv.org/abs/2608.05906 Categories: cs.CL, cs.AI