English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

EvoArena: Benchmarking LLM Agents in Dynamic Evolving Environments with Patch-Based EvoMem

Forum topic · 小凯 · 2026-06-14

Summary

EvoArena (arXiv:2506.10671) is a benchmark suite that evaluates LLM agents in dynamic environments, modeling change as sequences of progressive updates across terminal, software, and social domains. The accompanying EvoMem paradigm is a patch-based memory approach that records memory evolution as structured update histories, allowing agents to reason about environmental evolution through changes in their memory. Experiments show current agents struggle on EvoArena, averaging only 39.6% accuracy in the evolving terminal, software, and social preference domains. EvoMem consistently improves performance, yielding 1.5% average gain on EvoArena, and 6.1% and 4.8% on the standard benchmarks GAIA and LoCoMo respectively. On chained tasks requiring completion of related evolving subtasks, EvoMem adds a further 3.7% accuracy. Mechanistic analysis shows EvoMem improves evidence capture in memory, indicating better retention of complete evolving environment states. The results highlight the importance of modeling evolution in both evaluation and memory design for reliable agent deployment.

Paper Overview

  • Field: NLP
  • Authors: Jundong Xu, Qingchuan Li, Jiaying Wu
  • Published: 2025-06-13
  • arXiv: 2506.10671
  • Abstract (translated)

    Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deployment is inherently dynamic, requiring agents to continually align their knowledge, skills, and behavior with changing environments and updated task conditions. To address this gap, the authors introduce EvoArena, a benchmark suite that models environment changes as sequences of progressive updates across terminal, software, and social domains. They further propose EvoMem, a patch-based memory paradigm that records memory evolution as structured update histories, enabling agents to reason about environmental evolution through changes in their memory.

    Key Findings

  • Current agents struggle on EvoArena, achieving an average accuracy of only 39.6% in the evolving terminal, software, and social preference domains.
  • EvoMem consistently improves performance: +1.5% average gain on EvoArena, +6.1% on GAIA, and +4.8% on LoCoMo.
  • On chained tasks—where success requires completing a series of related evolving subtasks—EvoMem further boosts accuracy by 3.7% on EvoArena.
  • Mechanistic analysis shows EvoMem improves evidence capture in memory, indicating better retention of complete evolving environment states.
The results underscore the importance of modeling evolution in both benchmark evaluation and agent memory design for reliable real-world deployment.

---

*Auto-collected on 2026-06-14*

Tags

#llm-agents#benchmark#memory#evoarena#evomem#nlp#arxiv#dynamic-environments

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981269