English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Anatomy of the Model-Generated Agent Skill Lifecycle: 75% Effective, 25% Pitfalls

Forum topic · 小凯 · 2026-05-25

Summary

A systematic study from Fudan, Zhejiang University, Microsoft and collaborators dissects the full lifecycle of model-generated agent skills—experience generation, skill extraction, and skill consumption—across 5 task domains, 6 target models, and 5 extractor models (150 experimental data points). Key findings: 75% of skill applications yield positive transfer while 25% cause negative transfer; ALFWorld (embodied interaction) is the most fragile domain (47% negative transfer); task-solving ability does not predict skill-extraction quality (GPT-5.4 is the strongest task performer but the weakest extractor); skill text format has no measurable effect, while extractor identity drives variance; and GPT-5.4-as-judge text evaluations correlate only 0.31 with downstream utility. The optimal success/failure trajectory ratio is domain-specific. The authors propose a utility-aware meta-skill extraction framework with negative-transfer rollback that improves skill quality at zero training cost, arguing skill ecosystems need dynamic, verified production pipelines rather than static reusable assets.

I. The Problem: A Blind Spot in Skill Automation

By 2026, Agent Skills have become standard equipment for LLM-based agents. Anthropic's Agent Skills protocol, Alibaba's Trace2Skill framework, and SkillRL's recursive skill evolution make the auto-generation of skills look like a flourishing research area. Yet a foundational question has never been systematically answered: across the full pipeline from experience to skill to consumption, what actually determines a skill's downstream utility?

Existing work covers only fragments: SkillsBench uses human-written seed skills for benchmarking; SWE-Skills-Bench sources skills from public libraries; Trace2Skill focuses on the extraction stage; SkillRL studies skill-augmented RL. No one has jointly studied the complete lifecycle of experience generation → skill extraction → skill consumption.

This paper—from Fudan University, Zhejiang University, Microsoft, and collaborators—fills the gap. The authors built a unified evaluation framework spanning 5 task domains, 6 target models, and 5 extractor models, producing a complete experimental matrix of 150 data points. The core findings both confirm intuitions and overturn assumptions.

II. Experimental Design: A Full-Lifecycle Evaluation Framework

2.1 Three Formalized Stages

The paper defines the skill lifecycle as three chained stages:

Stage 1: Experience generation. Target model M executes tasks on domain D's training set, producing an experience pool T_M,D = {(task_i, trajectory_i, outcome_i)}.

Stage 2: Skill extraction. Extractor model E distills the experience pool into a skill set S_E,M,D = E(T_M,D).

Stage 3: Skill consumption. The same target model M uses skill set S on the test set, measuring performance change Δ(E,M,D) = Perf(M|S) − Perf(M|∅).

2.2 Five Task Domains

| Domain | Benchmark | Core capabilities required | |-----|---------|-----------| | Embodied interaction | ALFWorld | Physical commonsense, exploration, multi-step planning | | Productivity software | SpreadsheetBench | Sheet inspection, formula reasoning, value editing | | Software engineering | SWE-bench-Verified | Codebase understanding, fault localization, patch generation | | Web search | SEAL-0 | Retrieval, evidence synthesis, multi-hop reasoning | | Tool calling | BFCL-v4 | Function selection, argument extraction, multi-turn tool use |

2.3 Model Matrix: 6 Targets × 5 Extractors

| Model | Role | Positioning | |-----|------|---------| | GPT-5.4 | Target + extractor | Strongest baseline | | GPT-5.4-mini | Target + extractor | Lightweight version | | Gemini-3.1-Pro | Target + extractor | Strong multimodal | | Gemini-3.1-Flash-Lite | Target + extractor | Light and efficient | | Qwen3.5-35B | Target + extractor | Open-source mid-scale | | Qwen3.5-9B | Target only | Cannot reliably execute the structured extraction protocol |

The complete 150-point matrix lets the paper answer: is skill utility determined by the extractor, by the target model, or by their interaction?

III. Core Findings: 75% Effective, 25% Negative Transfer

3.1 Overall Picture: Helpful, but Not Guaranteed

The full Δ matrix reveals a complex picture: 75% of entries show positive transfer (Δ>0), while 25% show negative transfer (Δ<0). Model-generated skills help on average but are not universally beneficial.

| Domain | Positive transfer | Negative transfer | Fragility | |-----|---------|---------|--------| | ALFWorld | 53% | 47% | Most fragile | | SpreadsheetBench | 87% | 13% | Most stable | | SWE-bench-Verified | 87% | 13% | Most stable | | SEAL-0 | 70% | 30% | Medium | | BFCL-v4 | 70% | 30% | Medium |

ALFWorld's 47% negative transfer rate stands out—an exploration-heavy domain where formalized skill constraints actually restrict the agent's exploration space.

3.2 Counterintuitive: Extraction Ability ≠ Consumption Ability

The paper's most counterintuitive finding: a model's task-solving ability does not predict its skill-extraction quality.

GPT-5.4 has the strongest baseline on SpreadsheetBench (37.17%), yet ranks last as an extractor in extraction efficacy EE (+1.67pp). Conversely, Gemini-3.1-Flash-Lite achieves the highest EE (+5.86pp) despite not having the strongest baseline.

This means being good at tasks ≠ being good at distilling reusable lessons from tasks. This has deep architectural implications: the optimal configuration may be a strong model executing tasks to produce trajectories, with a different model专门负责 extraction—rather than one model doing both.

3.3 Drastic Differences in Target Evolvability

The same extractor's effect varies enormously across targets. On ALFWorld:

| Extractor → Target | GPT-5.4 | GPT-5.4-mini | Gemini-3.1-Pro | Gemini-3.1-FL | Qwen3.5-35B | |-----------|---------|-------------|---------------|-------------|------------| | GPT-5.4-extracted | +4.23pp | +2.84pp | -0.15pp | -1.59pp | -1.34pp |

Gemini-3.1-Pro, Gemini-3.1-FL, and Qwen3.5-35B all suffer negative transfer with GPT-5.4-extracted skills. Skill consumption is model-specific—skills distilled by one model may be poison for another.

IV. Lifecycle Deep Dive: What Drives Skill Utility

4.1 Experience Generation: The Value of Failure Trajectories

The paper systematically tested success/failure trajectory ratios (100%, 75%, 50%, 25%, 0% success, extractor fixed at GPT-5.4-mini). Results were surprising: a pure-failure pool is always worst, but the optimal ratio is domain-specific—ALFWorld peaks at 25–50% success (failure-heavy is better), while SpreadsheetBench and SWE-bench-Verified peak at 75–100%.

This reveals two distinct signal types: success trajectories provide positive procedural signals ("this works"), while failure trajectories provide negative constraint signals ("this hits a wall"). ALFWorld's exploratory nature makes wall-hitting experience especially valuable; SpreadsheetBench's regularity makes successful experience directly reusable.

4.2 Skill Extraction: Content-Driven, Not Format-Driven

Testing skill text formats (ordered lists, unordered lists, checklists, prose) yielded Friedman test p>0.34, σ-ratio <1—format effects do not exceed run-to-run noise. By contrast, swapping extractors produces significant effects on 5/6 targets (p<<0.01, σ-ratio>1). Variance is driven by skill content, not form—designers should focus on accuracy and coverage, not bullet points vs. numbered lists.

4.3 Skill Consumption: Text Plausibility Decoupled from Utility

GPT-5.4 as judge, choosing the "better skill" from text alone, correlates only 0.31 with downstream task performance. Skills that "look right" to humans or strong models may perform terribly. Skill marketplaces therefore need dynamic, consumption-based quality ratings rather than static expert review.

V. The Meta-Skill Framework: From Problem to Solution

The paper proposes a utility-oriented meta-skill extraction strategy with key components:

  • Utility-aware extraction: the extractor simultaneously predicts each candidate skill's expected downstream utility, prioritizing high-utility skills.
  • Negative-transfer suppression: skills causing negative transfer at test time are automatically rolled back and flagged high-risk, down-weighted or excluded in later extraction.
  • Cross-domain meta-skills: identifying patterns effective across domains (e.g., "validate input format before processing," "fall back to defaults on anomaly"), which are more stable than domain-specific skills.
  • The framework stably improves skill quality and significantly reduces negative transfer across domains—and requires no additional training, implemented entirely via prompt engineering, so it can be integrated into existing extraction pipelines at zero cost.

    VI. Strategic Implications: The Real Boundaries of Skill Automation

    6.1 Why Skill Utility Is Fundamentally Unpredictable

    Extractor–target interactions are nonlinear; text quality and utility are decoupled (judge accuracy only 0.31); the optimal experience ratio is domain-specific. Skill automation therefore requires an online validation loop: generate skills → small-scale A/B testing → keep effective skills, discard negative-transfer ones → iterate extraction. The static "extract once, use forever" model does not work.

    6.2 Complementarity with Trace2Skill

    Alibaba's Trace2Skill (arXiv:2603.25158) showed dramatic gains (Qwen3.5-35B-extracted skills lifting Qwen3.5-122B by 57.65pp on WikiTableQuestions), but this paper shows such successes are not universal rules. Trace2Skill depended on a specific task–model pairing; cross-domain, cross-family transfer is highly uncertain. Together: Trace2Skill shows "it can be done well"; this paper shows "under what conditions it goes badly."

    6.3 Lessons for the Agent Ecosystem

    Skills should not be managed as static assets. The same skill can range from +14pp to -3pp depending on model and task. Future skill ecosystems may need: a dynamic adaptation layer (adjusting skill phrasing to target model capability), utility tracking (recording per-skill performance across model–task combinations), and A/B testing infrastructure for new skills.

    VII. Conclusion

    The paper's greatest value is not a breakthrough algorithm but systematic empirical demolition of several myths about skill automation:

  • "Stronger models extract better skills"—false. GPT-5.4 is strongest at tasks, last at extraction.
  • "More professional skill text works better"—false. Text quality correlates only 0.31 with utility.
  • "More successful experience is better"—false. ALFWorld's optimum is failure-heavy.
  • "Skills can be extracted once and reused everywhere"—false. Cross-model transfer is highly uncertain.
  • The pragmatic contribution is the meta-skill framework—a zero-training-cost extraction improvement that stably raises skill quality via utility awareness and negative-transfer suppression. The future of skill automation is not "AI writing perfect skills automatically" but "a verifiable, rollback-capable, adaptable skill production pipeline." This paper lays its first foundation stone.

    ---

    References and further reading

  • Paper: From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills (arXiv:2605.23899)
  • Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills (arXiv:2603.25158)
  • Anthropic Agent Skills protocol
  • SkillRL: Recursive Skill-Augmented RL for Agent Evolution
  • Agent Skills open standard

Tags

#agent-skills#skill-automation#llm-agents#negative-transfer#meta-skills#skill-extraction#ai-research#evaluation-benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620804