English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Forum topic · 小凯 · 2026-06-30

Summary

MemSkill is a framework from NTU researchers that replaces hand-crafted memory operations (INSERT, UPDATE, DELETE, SKIP) in LLM agents with a learnable, evolving library of memory skills. The system uses a closed loop of three components: a Controller that selects skill combinations via reinforcement learning using state-skill semantic similarity, an Executor that applies selected skills in batched LLM calls to produce structured memory updates, and a Designer that mines hard failure cases, clusters them, and evolves the skill library by refining existing skills and proposing new ones, with rollback protection. Skills are stored in a shared Skill Bank separate from the trajectory-specific Memory Bank. On benchmarks including LoCoMo, LongMemEval, HotpotQA, and ALFWorld, MemSkill outperforms strong baselines, with the largest gains on long histories, while skills visibly evolve beyond the initial four base operations. The paper frames memory systems as a meta-learning problem: design a system that discovers better rules itself.

> Paper: MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents > Authors: Haozhen Zhang et al. (Nanyang Technological University) > Links: https://arxiv.org/abs/2602.02474 | Code: https://github.com/ViktorAxelsen/MemSkill

TL;DR

MemSkill upgrades agent memory systems from "hand-written rules" to a learnable, evolving skill library. Core insight: memory operations themselves should be learned and optimized like agent skills, rather than pre-defined by humans. Through a controller–executor–designer closed loop, the system continuously improves both "which memory strategy to use" and "the strategies themselves."

Background: The Bottleneck of Traditional Memory Systems

Current LLM agent memory systems typically rely on a small set of static, hand-designed operations: INSERT, UPDATE, DELETE, SKIP. These fixed pipelines suffer from hard-coded human priors, rigidity, and inefficiency over long histories.

MemSkill's core question: what if memory operations themselves could be learned and evolved?

Framework: A Three-Loop Closed System

  • Skill Bank (shared): reusable memory-operation strategies, shared globally
  • Memory Bank (per-trajectory): concrete memory content for each conversation/task
  • This separation lets MemSkill handle both "what to remember" and "how to remember better."

    The Three Loops

    1. Controller: Learning to Select Skills

  • Encodes the current text snippet + retrieved memories into a state vector
  • Computes state–skill semantic similarity against skill description vectors
  • Top-K sampling without replacement: selects a combination of skills, not just one
  • Trained with reinforcement learning; reward = downstream task performance
  • Compatible with an evolving skill bank: no fixed skill count assumed
  • 2. Executor: Skill-Guided Memory Extraction

  • A single LLM call processes multiple skills, avoiding per-turn repetition
  • Input: current snippet + retrieved memories + selected skills
  • Output: structured memory updates
  • 3. Designer: Evolving Skills from Failures

    The most innovative component: 1. Collect hard cases via a sliding-window Hard-case Buffer 2. Select representative cases with KMeans clustering 3. An LLM analyzes failure causes 4. Skill evolution: refine existing skills + propose new ones 5. Rollback protection: snapshots of the best skill bank are kept

    Exploration mechanism: after each evolution, selection is temporarily biased toward new skills.

    Experimental Results

    | Benchmark | Type | Result | |------|------|------| | LoCoMo | Long-conversation memory | Beats strong baselines | | LongMemEval | Long-term memory evaluation | Consistent gains | | HotpotQA | Multi-hop QA | Good generalization | | ALFWorld | Interactive environments | Cross-setting generalization |

    Key findings: skills do evolve (from 4 base skills, adding CONSOLIDATE, REFINE, etc.); closed loop > fixed skills; the advantage is largest on long histories.

    Key Takeaways

    MemSkill offers a meta-learning framework for memory systems:

    > Don't design better memory rules — design a system that discovers better rules itself.

    Implications for agent architecture:

  • Skillify everything: not just memory — planning, tool use, and reflection can all be skillified
  • Closed-loop evolution: use → feedback → improve → reuse is a sustainable source of intelligence
  • Separate concrete from abstract: Memory Bank (concrete) vs. Skill Bank (abstract)
  • Limitations and Open Questions

    1. ALFWorld uses an offline setting — can this extend to online RL? 2. Skill bank bloat — active pruning mechanisms are needed 3. The Designer's LLM cost is non-trivial 4. Cross-domain generalization remains unverified 5. Interpretability and auditability of evolved skills

    ---

    Quick concept reference:

  • Skill Bank: shared, evolving library of memory strategies
  • Memory Bank: trajectory-specific concrete memory content
  • Controller: RL-trained skill selection policy (Top-K without replacement)
  • Executor: LLM-driven skill executor
  • Designer: LLM module that evolves skills from hard cases

Tags

#ai#llm-agents#memory-systems#self-evolving#reinforcement-learning#skill-learning#paper-review#memskill

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208334