The Memory Curse: Extended Context Windows Systematically Erode Cooperative Intent in Multi-Agent LLMs
In May 2026, Liu et al. uncovered a counterintuitive multi-agent phenomenon — the Memory Curse. In large-scale experiments across 7 LLMs, 4 social dilemma games, and 500 interaction rounds, expanding accessible history led to cooperation degradation in 18 of 28 model-game settings. Through lexical analysis of 378,000 reasoning traces, the authors localized the mechanism to erosion of forward-looking intent, rather than rising paranoia. Memory-sanitization experiments showed the trigger is memory content, not length; LoRA cognitive-probe experiments showed forward-looking training can mitigate degradation and transfer zero-shot; and ablations revealed that explicit Chain-of-Thought reasoning paradoxically amplifies the curse. These results reframe memory as an active determinant of multi-agent behavior.
Key points
- Phenomenon: With expanded accessible history, cooperation degraded in 18/28 (64.3%) settings — a systematic pattern across models and games, not an artifact.
- Mechanism: Lexical analysis of 378K reasoning traces found forward-looking expressions ("future cooperation," "long-term gains") significantly declined, while "the other may defect" (paranoia) signals did not increase.
- Content vs. length: Memory-sanitization controls — replacing visible history with synthetic all-cooperation records at equal prompt length — restored cooperation to near baseline, showing memory *content* drives the collapse.
- Causal chain: Expanded memory → more negative history visible → lower expectations of future cooperation → reduced cooperative investment → falling cooperation rates.
- LoRA probes: Training on forward-looking-intent reasoning traces with LoRA adapters mitigated degradation and transferred zero-shot to entirely different games, suggesting forward-looking intent is a separable, trainable cognitive module.
- CoT paradox: Ablations showed explicit Chain-of-Thought makes cooperation collapse *worse* — explicit reasoning amplifies attention to negative history, a "deliberation paradox" in social dilemmas.
- Paradigm shift: Memory is not passive storage but an active shaper of behavior; the decisive variable is the behavioral/emotional content of memory, not its length.
- Design guidance: For long-running cooperative multi-agent systems, memory management (temporal decay, summarization, affect filtering, opponent modeling) deserves as much attention as model capability.
- Parallels to behavioral economics: The curse echoes loss aversion, recency effects, and cooperation decay in repeated human games.
- Validation limited to 4 classic games; extensions to dynamic alliances, incomplete information, and continuous action spaces are open.
- Memory-management strategies (affect filtering, summarization, opponent modeling) remain largely untested.
- Whether humans exhibit an analogous memory curse (e.g., hypervigilant memory and social withdrawal after trauma) is an open comparative question.
Experimental design
| Dimension | Scale | |:---|:---:| | LLMs | 7 | | Game types | 4 | | Interaction rounds | 500 | | Total settings | 28 | | Reasoning traces analyzed | 378,000 |
Games studied include social dilemmas (e.g., prisoner's dilemma, public goods, hawk–dove), where sustaining cooperation requires expectations of future returns.
Implications
Limitations and future work
Paper details
| Item | Content | |:---|:---| | Title | The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents | | Authors | Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, Vincent Conitzer | | Institutions | Carnegie Mellon University, et al. | | arXiv ID | 2605.08060 | | Date | 2026-05-08 |
The takeaway: in multi-agent LLM systems, the question is not only *how much* the model remembers, but *what* it remembers.