English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

Forum topic · 小凯 · 2026-05-26

Summary

This arXiv paper (2505.21422) by Zisu Huang, Jingwen Xu, and Yifan Yang presents the first comprehensive study of the full skill lifecycle in language agents: experience generation, skill extraction, and skill consumption. The authors build a utility-grounded evaluation framework and run systematic experiments across extractors and consumer agents over five diverse agent task domains. Key findings: model-generated skills are beneficial on average but exhibit non-trivial negative transfer, and neither extractor nor consumer behavior is consistent—a model can be a strong skill extractor yet a weak consumer, and vice versa. Skill utility does not correlate with model scale or baseline task strength. The paper analyzes how experience composition shapes skill quality, what characterizes useful skills, and how the same skill transfers across different consumers. Based on these insights, the authors derive a concrete meta-skill that steers skill extraction toward features correlated with actual utility, consistently improving skill quality across domains and substantially reducing negative transfer. Compiled from zhichai.net, originally published 2026-05-26.

Paper Overview

Research area: Machine Learning Authors: Zisu Huang, Jingwen Xu, Yifan Yang Published: 2026-05-26 arXiv: 2505.21422

Abstract

Language agents increasingly improve by reusing skills — structured procedural artifacts distilled from past experience. In particular, domain-level and model-generated skills are especially promising: they offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle — experience generation, skill extraction, and skill consumption — to ask whether such skills actually work, when they work, and what makes them succeed or fail.

Key Findings

  • Utility-grounded evaluation framework: The authors systematically evaluate skills across extractors and consumer agents, spanning five diverse agent task domains.
  • On average beneficial, but: Model-generated skills help on average, yet exhibit non-trivial negative transfer.
  • Inconsistent roles: Neither extractor nor consumer behavior is consistent — a model can be a strong skill extractor yet a weak consumer, or vice versa.
  • No scale shortcut: Skill utility is uncorrelated with model size or baseline task strength.
  • Lifecycle analysis: The paper examines how experience composition shapes skill quality, what features characterize useful skills, and how the same skill transfers across different consumers.
  • Meta-skill contribution: The findings are distilled into a concrete "meta-skill" that guides skill extraction toward features correlated with actual utility, consistently improving skill quality across domains and substantially reducing negative transfer.

Original Abstract (excerpt)

> Language agents increasingly improve by reusing skills — structured procedural artifacts distilled from past experience. In particular, domain-level and model-generated skills are especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle — experience generation, skill extraction, and skill consumption — to ask whether such skills actually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and...

---

*Auto-collected on 2026-05-26 via zhichai.net*

Tags

#machine-learning#llm-agents#skill-learning#arxiv#agent-frameworks#evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620812