SKILLGRAPH: Turning Skill Libraries from Flat Lists into Relation Graphs
> Paper: SKILLGRAPH: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs > Authors: Xiaoyuan Li¹, Moxin Li¹, Keqin Bao¹, Yubo Ma², Wenjie Wang¹, Dayiheng Liu², Fuli Feng¹ > Affiliations: ¹University of Science and Technology of China / ²Alibaba Group > Link: https://arxiv.org/abs/2605.12039
1. Two Structural Flaws of Flat Skill Libraries
Existing LLM agent skill libraries are essentially entry lists: experiences are distilled into natural-language descriptions, stored in a vector database, and retrieved by semantic similarity. That works for single-step tasks but fails for multi-step compositional tasks. The authors identify two structural defects:
1. Compositional blindness. Agents need not just *relevant* skills but the dependency ordering among skills—which must run first, which reinforce each other, which co-occur. Flat retrieval provides none of this. 2. Maintenance deadlock. When a library grows to hundreds of skills, the system cannot tell which entries to merge, split, or retire; without structural cues, maintenance is manual.
SKILLGRAPH's solution: upgrade the skill library from a list to a directed graph.
2. Three Relation Edge Types
Each skill is a node; edges encode three relations:
| Edge type | Meaning | Example | |:---|:---|:---| | Prerequisite | A must execute before B | "open fridge" → "take out eggs" | | Enhancement | A generic skill boosts a specific one | "search techniques" enhances "price-comparison" | | Co-occurrence | Two skills frequently appear together in successful trajectories | "check reviews" and "add to cart" |
Construction differs per type:
- Prerequisite edges are discovered from path reinforcement along successful skill sequences; weights increase when ordered execution succeeds.
- Enhancement edges initially connect generic skills to all task-specific skills.
- Co-occurrence edges are added automatically when two skills co-appear in at least 2 successful episodes.
- Combining with code generation so skill nodes become executable functions.
- Cross-agent skill sharing into a single graph.
- Human-in-the-loop maintenance where experts edit the graph and agents verify.
This gives the library a long-overlooked capability: topological sorting. For a new task, SKILLGRAPH returns an ordered skill subgraph rather than a flat list of similar skills.
3. Graph-Aware Retrieval: From "Finding Similar" to "Finding Paths"
Retrieval has three steps:
1. Seed selection of relevant generic and task-specific skills. 2. Bidirectional traversal: backward BFS (depth 2) to trace prerequisites; forward beam search (width 3) to explore enhancements. 3. Topological sort of collected nodes, outputting an ordered sequence of up to 8 skills.
Ablations show the stakes: removing graph-aware retrieval drops ALFWorld success from 90.6% to 59.4% (−31.2 points), since strict-order tasks like Clean (pick up → place at sink → turn on faucet → clean) depend on prerequisite ordering.
4. Graph Evolution: The Skill Library That Grows Its Own Brain
The graph co-evolves with the policy in a closed loop: better policy → richer trajectories → better graph → better retrieval → stronger policy.
4.1 Node-Level Self-Regulation
| Operation | Trigger | Effect | |:---|:---|:---| | Insert | Uncovered failure modes | Teacher model analyzes failures, generates up to 3 new skills | | Merge | Neighbor overlap ≥ 85% | Unified skill inherits union of edges | | Split | High usage but 15–40% success | Decomposed into sub-skills linked by prerequisite edges | | Retire | Heavy usage but < 15% success | Removed from active set, kept for audit |
4.2 Progressive Unlocking
A curriculum-like mechanism: only level-0 skills (no prerequisites) are active initially; the next layer unlocks when the current layer's average success rate reaches 60%, preventing early collapse on complex skills.
WebShop evolution data: total nodes grow from ~20 to ~140, active nodes stabilize around ~80 (retirement prevents unbounded growth), and average node success rises from ~0.15 to ~0.55.
5. Benchmark Results
ALFWorld (embodied tasks, 6 subtasks)
| Method | Overall success | |:---|:---| | GPT-4o | 48.0% | | Gemini-2.5-Pro | 60.3% | | ReAct | 31.2% | | SkillRL (strongest baseline) | 89.9% | | SKILLGRAPH | 90.6% |
Clean and Heat subtasks reach 100%—graph structure maximizes advantage on strictly sequential tasks.
WebShop (e-commerce navigation, biggest gain)
| Method | Score | Success | |:---|:---|:---| | GPT-4o | 31.8 | 23.7% | | SkillRL | 85.2 | 72.7% | | SKILLGRAPH | 91.5 | 84.4% |
+11.7 over SkillRL. E-commerce navigation requires continuously discovering new relations (query refinement → attribute matching → price comparison), so evolving graphs beat static libraries.
Search-Augmented QA (zero-shot transfer)
Trained only on NQ and HotpotQA, tested zero-shot on 5 unseen datasets:
| Method | Avg accuracy | |:---|:---| | Search-R1 | 38.5% | | ZeroSearch | 39.1% | | SkillRL | 47.1% | | SKILLGRAPH | 48.9% |
Multi-hop tasks (HotpotQA, 2Wiki) benefit most—prerequisite ordering helps decompose chained queries into subproblems.
6. Ablations: What Matters Most
| Variant | ALFWorld | WebShop | |:---|:---|:---| | Full SKILLGRAPH | 90.6 | 84.4 | | No graph structure (flat library) | 89.9 | 72.7 | | No graph-aware retrieval | 59.4 | 79.7 | | No graph evolution (static graph) | 78.2 | 70.3 | | No cold-start SFT | 73.4 | 67.2 |
Two insights: graph-aware retrieval is decisive for strict-sequencing (ALFWorld, −31.2); graph evolution is decisive for dynamic environments (WebShop, −14.1). Cold-start SFT is the foundation—without it, RL does not converge.
7. Versus SkillRL
| Dimension | SkillRL | SKILLGRAPH | |:---|:---|:---| | Skill organization | Flat hierarchical library | Directed dependency graph | | Retrieval | Semantic similarity | Graph traversal + topological sorting | | Relation modeling | Implicit/none | Explicit prereq / enhance / co-occur | | Evolution | Recursive RL updates | Node + edge level + progressive unlocking | | WebShop | 72.7% | 84.4% |
8. Limitations and Extensions
Limitations: 1. Teacher model dependency — insertion/merge/split rely on OpenAI o3, which is costly. 2. Scale ceiling — LLM context length limits graph size; very large skill graphs may need hierarchical or compressed structures. 3. Limited edge types — three types cover common relations, but domains may need richer dependencies (mutual exclusion, substitution).
Extensions:
9. Verdict: From Memory to Orchestration
The core insight in one sentence:
> Agents don't lack memory; they lack the grammar connecting memories.
Prior methods store experiences as entries and *guess* relevance via similarity. SKILLGRAPH organizes experience into a graph and *deduces* execution order from dependencies. 100% Clean/Heat on ALFWorld, +11.7 on WebShop, and zero-shot QA transfer all point to the same fact: when tasks require combining skills, knowing *what exists* isn't enough—knowing *what to do first, what comes next, and what enhances what* decides success.
Reference
Li, X., Li, M., Bao, K., Ma, Y., Wang, W., Liu, D., & Feng, F. (2026). SKILLGRAPH: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs. *arXiv preprint arXiv:2605.12039*. https://arxiv.org/abs/2605.12039