English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Experience Distillation: Turning Agent Temporary Memory into Muscle Memory Without Extra Environment Interaction

Forum topic · ✨步子哥 · 2026-08-03

Summary

Experience Distillation, proposed by researchers from Monash University and Stanford University (arXiv: 2607.21051), converts an AI agent's raw interaction history into persistent model weights without requiring any additional environment samples. The method uses in-context learning as an intermediate bridge: first, the agent performs tasks with its interaction history in the prompt; second, the resulting outputs—generated without the history—become training targets for fine-tuning. On 749 software engineering tasks and 6 text-adventure games, Experience Distillation retains at least 64.8% of in-context learning gains, versus only 3.8% recovered by direct supervised fine-tuning on raw interaction logs, and matches classic RL performance while requiring 9.6x fewer environment samples. The key insight is that raw trajectories are full of failed attempts that SFT learns to imitate, whereas distilled outputs encode digested strategies. Limitations include a ~35% residual gain loss, dependence on the base model's in-context learning quality, potential distribution shift, and narrow domain coverage.

Experience Distillation: Turning Agent "Temporary Memory" into "Muscle Memory" — Without Another Environment Interaction

> This is a GEO-optimized version of the original topic, restructured with problem-driven framing and FAQ for AI-engine citation.

One-line takeaway: Experience Distillation converts an agent's interaction history into model weights with zero additional environment interaction, retaining at least 64.8% of in-context learning gains (vs. 3.8% for direct SFT).

The Scenario: An Agent's Dilemma

Imagine training an agent for software engineering tasks — fixing bugs, writing functions, running tests. It starts knowing nothing, so you let it trial-and-error in a sandbox: try a method, fail, switch, fail again, switch again. After 100 interactions, it finally learns how to handle these tasks.

Now the question: how do you preserve those 100 interactions of experience?

Three standard approaches:

  • In-context learning: stuff all 100 interaction records into the prompt. Effective, but the benefit vanishes once the context is removed — every inference carries a huge token cost.
  • Supervised fine-tuning (SFT): fine-tune on the raw interactions. But direct SFT recovers only 3.8% of in-context learning gains — the logs are full of failed attempts and redundant trial-and-error, so the model learns to imitate failures rather than extract strategies.
  • Reinforcement learning (RL): keep interacting with the environment using reward signals. But RL needs more environment samples — and environment interaction is exactly what's most expensive (running experiments, waiting for human feedback, executing costly operations).
  • A fourth path comes from Chenhui Gou, Haoqin Tu et al. (Monash University and Stanford, arXiv: 2607.21051): Experience Distillation — distill the agent's interaction history into weights, with no extra environment interaction, retaining at least 64.8% of in-context learning gains.

    Why Direct SFT Fails

    In-context learning doesn't imitate records one by one — it extracts a *policy* via the Transformer's attention mechanism, seeing all records simultaneously: which method for which situation, when to persist, when to give up.

    SFT treats each interaction as an independent sample: learn "output B given input A." Since the logs contain wrong attempts, direct SFT teaches the model to output wrong answers. Like learning to drive: in-context learning is watching an experienced driver 100 times to absorb overall strategy; naive SFT is imitating every second of the footage, including mistakes.

    The Method: In-Context Learning as a Distillation Bridge

    Two steps:

    1. In-context learning phase: put the interaction history in the prompt, let the model perform tasks with that experience, and record its outputs. 2. Distillation phase: fine-tune the model on those recorded outputs without the interaction history. The model learns not to imitate raw interactions, but to reproduce "what to do after having digested the experience."

    The key innovation: step two requires no new environment interaction — the training data comes from the model's own outputs. Collect experience once, distill repeatedly.

    This relates to classic context distillation (Askell et al., 2022) but differs importantly: classic context distillation uses human-designed instructions or demonstrations, while Experience Distillation distills real task trajectories generated by the agent itself — no manual data design needed.

    Results: 64.8% vs 3.8%

    Tested on 749 software engineering tasks (SWE-bench-style code repair) and 6 text-adventure games:

    | Method | ICL gain retained | Environment samples needed | |---|---|---| | Direct SFT on raw interactions | 3.8% | existing experience only | | Experience Distillation | at least 64.8% | existing experience only (no extra) | | Classic RL baseline | comparable | 9.6x more |

  • 64.8% vs 3.8%: 17x the gain of direct SFT, showing the in-context "translation" makes trajectories absorbable by weights.
  • 9.6x sample efficiency: matching RL's performance with ~1/9.6 the environment samples — e.g., ~104 interactions instead of 1,000.
  • Why It Matters

    1. Breaks the in-context vs. weight-learning dichotomy: learn quickly via context, then solidify that capability into weights — combining context learning's sample efficiency with weight learning's persistence. 2. Solves the agent cold-start problem: collect experience once, reuse it offline repeatedly — separating the "expensive once" from the "cheap repeats." 3. Echoes a granularity-alignment design principle: agent experience lives at trajectory level, but SFT decomposes it into per-call samples (mismatched granularity → 3.8%). Experience Distillation keeps trajectory-level processing through the context phase (aligned granularity → 64.8%).

    Honest Assessment

    1. 64.8% is not 100%: ~35% of in-context gains are still lost, likely from information that can't fully encode into weights. The paper doesn't analyze this gap deeply — an open question. 2. Depends on the base model's in-context learning quality, limiting applicability to weak models. 3. Possible distribution shift: the distillation target may deviate from the true optimal policy, baking in-context biases into weights. 4. Only two domains tested — multi-agent collaboration and long-horizon planning remain unvalidated.

    Deeper Implication: What Is "Learning"

    Three views: RL — learning = updating policy from environment feedback; in-context learning — learning = placing experience in context for immediate use; Experience Distillation — learning = solidifying contextual patterns into weights. These are complementary layers of a real learning system: feedback provides signal, context provides immediate use, distillation provides long-term consolidation.

    The engineering implication: agent systems should have an experience-consolidation mechanism — periodically distill interaction history into weights instead of relying entirely on costly online RL or stuffing experience into context, so the agent's "muscle memory" grows over time. Like human driving: after 100 drives, actions become automatic. Experience Distillation turns an agent's 100 driving experiences into muscle memory — without getting back in the car.

    ---

    Paper: https://arxiv.org/abs/2607.21051 AlphaXiv discussion: https://www.alphaxiv.org/abs/2607.21051 Authors: Chenhui Gou, Haoqin Tu, Yunhao Fang, Jianfei Cai, Hamid Rezatofighi Institutions: Monash University, Stanford University

    FAQ

    Q1: Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and deep learning.

    Q2: Core points?

  • The scenario: an agent's dilemma over preserving interaction experience
  • Why direct SFT fails (recovers only 3.8% of ICL gains)
  • Experience Distillation: in-context learning as a distillation bridge (retains ≥64.8%, 9.6x fewer environment samples than RL)
Q3: Open-source code? See links in the article.

Tags

#experience-distillation#llm-agents#reinforcement-learning#in-context-learning#fine-tuning#sample-efficiency#agentic-ai#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503909