Summary
EvoArena (arXiv:2506.10671, by Jundong Xu, Qingchuan Li, and Jiaying Wu) is a benchmark suite that evaluates LLM agents in dynamic rather than static environments. It models environment changes as sequences of progressive updates across terminal, software, and social domains. The authors also propose EvoMem, a patch-based memory paradigm that records memory evolution as structured update histories, letting agents reason about environmental evolution through memory changes. Experiments show current agents struggle on EvoArena, averaging only 39.6% accuracy across the evolving domains. EvoMem consistently improves performance, with average gains of 1.5% on EvoArena, 6.1% on GAIA, and 4.8% on LoCoMo. On chained tasks, EvoMem further boosts accuracy by 3.7%, where success requires completing a series of related evolving subtasks. Mechanistic analysis shows EvoMem improves evidence capture in memory, preserving a more complete evolving environmental state. The results highlight the importance of modeling evolution in both evaluation and memory for reliable agent deployment.
Paper Overview
- Research Area: NLP
- Authors: Jundong Xu, Qingchuan Li, Jiaying Wu
- Published: 2025-06-13
- arXiv: 2506.10671
Abstract
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deployment is inherently dynamic, requiring agents to continually align their knowledge, skills, and behavior with changing environments and updated task conditions.
To address this gap, the authors introduce EvoArena, a benchmark suite that models environment changes as sequences of progressive updates across terminal, software, and social domains. They further propose EvoMem, a patch-based memory paradigm that records memory evolution as structured update histories, enabling agents to reason about environmental evolution through changes in their memory.
Key Findings
- Current agents struggle on EvoArena, achieving an average accuracy of only 39.6% across the evolving terminal, software, and social preference domains.
- EvoMem consistently improves performance: an average gain of 1.5% on EvoArena, 6.1% on GAIA, and 4.8% on LoCoMo.
- On chained tasks, EvoMem further boosts accuracy by 3.7% on EvoArena — success requires completing a series of related evolving subtasks.
- Mechanistic analysis shows EvoMem improves evidence capture in memory, indicating better preservation of the complete evolving environmental state.
The results underscore the importance of modeling evolution in both evaluation and memory for reliable agent deployment.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981192