> Paper: MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents > Authors: Haozhen Zhang et al. (Nanyang Technological University) > Links: https://arxiv.org/abs/2602.02474 | Code: https://github.com/ViktorAxelsen/MemSkill
TL;DR
MemSkill upgrades agent memory systems from "hand-written rules" to a learnable, evolving skill library. Core insight: memory operations themselves should be learned and optimized like agent skills, rather than pre-defined by humans. Through a controller–executor–designer closed loop, the system continuously improves both "which memory strategy to use" and "the strategies themselves."
Background: The Bottleneck of Traditional Memory Systems
Current LLM agent memory systems typically rely on a small set of static, hand-designed operations: INSERT, UPDATE, DELETE, SKIP. These fixed pipelines suffer from hard-coded human priors, rigidity, and inefficiency over long histories.
MemSkill's core question: what if memory operations themselves could be learned and evolved?
Framework: A Three-Loop Closed System
- Skill Bank (shared): reusable memory-operation strategies, shared globally
- Memory Bank (per-trajectory): concrete memory content for each conversation/task
- Encodes the current text snippet + retrieved memories into a state vector
- Computes state–skill semantic similarity against skill description vectors
- Top-K sampling without replacement: selects a combination of skills, not just one
- Trained with reinforcement learning; reward = downstream task performance
- Compatible with an evolving skill bank: no fixed skill count assumed
- A single LLM call processes multiple skills, avoiding per-turn repetition
- Input: current snippet + retrieved memories + selected skills
- Output: structured memory updates
- Skillify everything: not just memory — planning, tool use, and reflection can all be skillified
- Closed-loop evolution: use → feedback → improve → reuse is a sustainable source of intelligence
- Separate concrete from abstract: Memory Bank (concrete) vs. Skill Bank (abstract)
- Skill Bank: shared, evolving library of memory strategies
- Memory Bank: trajectory-specific concrete memory content
- Controller: RL-trained skill selection policy (Top-K without replacement)
- Executor: LLM-driven skill executor
- Designer: LLM module that evolves skills from hard cases
This separation lets MemSkill handle both "what to remember" and "how to remember better."
The Three Loops
1. Controller: Learning to Select Skills
2. Executor: Skill-Guided Memory Extraction
3. Designer: Evolving Skills from Failures
The most innovative component: 1. Collect hard cases via a sliding-window Hard-case Buffer 2. Select representative cases with KMeans clustering 3. An LLM analyzes failure causes 4. Skill evolution: refine existing skills + propose new ones 5. Rollback protection: snapshots of the best skill bank are kept
Exploration mechanism: after each evolution, selection is temporarily biased toward new skills.
Experimental Results
| Benchmark | Type | Result | |------|------|------| | LoCoMo | Long-conversation memory | Beats strong baselines | | LongMemEval | Long-term memory evaluation | Consistent gains | | HotpotQA | Multi-hop QA | Good generalization | | ALFWorld | Interactive environments | Cross-setting generalization |
Key findings: skills do evolve (from 4 base skills, adding CONSOLIDATE, REFINE, etc.); closed loop > fixed skills; the advantage is largest on long histories.
Key Takeaways
MemSkill offers a meta-learning framework for memory systems:
> Don't design better memory rules — design a system that discovers better rules itself.
Implications for agent architecture:
Limitations and Open Questions
1. ALFWorld uses an offline setting — can this extend to online RL? 2. Skill bank bloat — active pruning mechanisms are needed 3. The Designer's LLM cost is non-trivial 4. Cross-domain generalization remains unverified 5. Interpretability and auditability of evolved skills
---
Quick concept reference: