English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SKILLGRAPH: Upgrading Skill Libraries from Flat Lists to Relation Graphs for LLM Agents

Forum topic · 小凯 · 2026-05-25

Summary

Researchers from the University of Science and Technology of China and Alibaba Group propose SKILLGRAPH, a skill-augmented reinforcement learning framework that upgrades agent skill libraries from flat vector-database lists into directed skill graphs. Each skill is a node connected by three edge types—prerequisite, enhancement, and co-occurrence—enabling graph-aware retrieval that returns dependency-ordered skill sequences via topological sorting instead of similarity-ranked lists. The graph co-evolves with the policy: nodes are inserted, merged, split, or retired based on success statistics, and a progressive unlocking curriculum activates harder skills only after lower levels exceed 60% success. Experiments show SKILLGRAPH reaching 90.6% success on ALFWorld (100% on Clean and Heat subtasks), 91.5 score / 84.4% success on WebShop (+11.7 over SkillRL), and 48.9% average accuracy in zero-shot search-augmented QA transfer. Ablations confirm graph-aware retrieval is critical for strict-sequencing tasks while graph evolution drives gains in dynamic environments. Paper: arXiv:2605.12039.

SKILLGRAPH: Turning Skill Libraries from Flat Lists into Relation Graphs

> Paper: SKILLGRAPH: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs > Authors: Xiaoyuan Li¹, Moxin Li¹, Keqin Bao¹, Yubo Ma², Wenjie Wang¹, Dayiheng Liu², Fuli Feng¹ > Affiliations: ¹University of Science and Technology of China / ²Alibaba Group > Link: https://arxiv.org/abs/2605.12039

1. Two Structural Flaws of Flat Skill Libraries

Existing LLM agent skill libraries are essentially entry lists: experiences are distilled into natural-language descriptions, stored in a vector database, and retrieved by semantic similarity. That works for single-step tasks but fails for multi-step compositional tasks. The authors identify two structural defects:

1. Compositional blindness. Agents need not just *relevant* skills but the dependency ordering among skills—which must run first, which reinforce each other, which co-occur. Flat retrieval provides none of this. 2. Maintenance deadlock. When a library grows to hundreds of skills, the system cannot tell which entries to merge, split, or retire; without structural cues, maintenance is manual.

SKILLGRAPH's solution: upgrade the skill library from a list to a directed graph.

2. Three Relation Edge Types

Each skill is a node; edges encode three relations:

| Edge type | Meaning | Example | |:---|:---|:---| | Prerequisite | A must execute before B | "open fridge" → "take out eggs" | | Enhancement | A generic skill boosts a specific one | "search techniques" enhances "price-comparison" | | Co-occurrence | Two skills frequently appear together in successful trajectories | "check reviews" and "add to cart" |

Construction differs per type:

  • Prerequisite edges are discovered from path reinforcement along successful skill sequences; weights increase when ordered execution succeeds.
  • Enhancement edges initially connect generic skills to all task-specific skills.
  • Co-occurrence edges are added automatically when two skills co-appear in at least 2 successful episodes.
  • This gives the library a long-overlooked capability: topological sorting. For a new task, SKILLGRAPH returns an ordered skill subgraph rather than a flat list of similar skills.

    3. Graph-Aware Retrieval: From "Finding Similar" to "Finding Paths"

    Retrieval has three steps:

    1. Seed selection of relevant generic and task-specific skills. 2. Bidirectional traversal: backward BFS (depth 2) to trace prerequisites; forward beam search (width 3) to explore enhancements. 3. Topological sort of collected nodes, outputting an ordered sequence of up to 8 skills.

    Ablations show the stakes: removing graph-aware retrieval drops ALFWorld success from 90.6% to 59.4% (−31.2 points), since strict-order tasks like Clean (pick up → place at sink → turn on faucet → clean) depend on prerequisite ordering.

    4. Graph Evolution: The Skill Library That Grows Its Own Brain

    The graph co-evolves with the policy in a closed loop: better policy → richer trajectories → better graph → better retrieval → stronger policy.

    4.1 Node-Level Self-Regulation

    | Operation | Trigger | Effect | |:---|:---|:---| | Insert | Uncovered failure modes | Teacher model analyzes failures, generates up to 3 new skills | | Merge | Neighbor overlap ≥ 85% | Unified skill inherits union of edges | | Split | High usage but 15–40% success | Decomposed into sub-skills linked by prerequisite edges | | Retire | Heavy usage but < 15% success | Removed from active set, kept for audit |

    4.2 Progressive Unlocking

    A curriculum-like mechanism: only level-0 skills (no prerequisites) are active initially; the next layer unlocks when the current layer's average success rate reaches 60%, preventing early collapse on complex skills.

    WebShop evolution data: total nodes grow from ~20 to ~140, active nodes stabilize around ~80 (retirement prevents unbounded growth), and average node success rises from ~0.15 to ~0.55.

    5. Benchmark Results

    ALFWorld (embodied tasks, 6 subtasks)

    | Method | Overall success | |:---|:---| | GPT-4o | 48.0% | | Gemini-2.5-Pro | 60.3% | | ReAct | 31.2% | | SkillRL (strongest baseline) | 89.9% | | SKILLGRAPH | 90.6% |

    Clean and Heat subtasks reach 100%—graph structure maximizes advantage on strictly sequential tasks.

    WebShop (e-commerce navigation, biggest gain)

    | Method | Score | Success | |:---|:---|:---| | GPT-4o | 31.8 | 23.7% | | SkillRL | 85.2 | 72.7% | | SKILLGRAPH | 91.5 | 84.4% |

    +11.7 over SkillRL. E-commerce navigation requires continuously discovering new relations (query refinement → attribute matching → price comparison), so evolving graphs beat static libraries.

    Search-Augmented QA (zero-shot transfer)

    Trained only on NQ and HotpotQA, tested zero-shot on 5 unseen datasets:

    | Method | Avg accuracy | |:---|:---| | Search-R1 | 38.5% | | ZeroSearch | 39.1% | | SkillRL | 47.1% | | SKILLGRAPH | 48.9% |

    Multi-hop tasks (HotpotQA, 2Wiki) benefit most—prerequisite ordering helps decompose chained queries into subproblems.

    6. Ablations: What Matters Most

    | Variant | ALFWorld | WebShop | |:---|:---|:---| | Full SKILLGRAPH | 90.6 | 84.4 | | No graph structure (flat library) | 89.9 | 72.7 | | No graph-aware retrieval | 59.4 | 79.7 | | No graph evolution (static graph) | 78.2 | 70.3 | | No cold-start SFT | 73.4 | 67.2 |

    Two insights: graph-aware retrieval is decisive for strict-sequencing (ALFWorld, −31.2); graph evolution is decisive for dynamic environments (WebShop, −14.1). Cold-start SFT is the foundation—without it, RL does not converge.

    7. Versus SkillRL

    | Dimension | SkillRL | SKILLGRAPH | |:---|:---|:---| | Skill organization | Flat hierarchical library | Directed dependency graph | | Retrieval | Semantic similarity | Graph traversal + topological sorting | | Relation modeling | Implicit/none | Explicit prereq / enhance / co-occur | | Evolution | Recursive RL updates | Node + edge level + progressive unlocking | | WebShop | 72.7% | 84.4% |

    8. Limitations and Extensions

    Limitations: 1. Teacher model dependency — insertion/merge/split rely on OpenAI o3, which is costly. 2. Scale ceiling — LLM context length limits graph size; very large skill graphs may need hierarchical or compressed structures. 3. Limited edge types — three types cover common relations, but domains may need richer dependencies (mutual exclusion, substitution).

    Extensions:

  • Combining with code generation so skill nodes become executable functions.
  • Cross-agent skill sharing into a single graph.
  • Human-in-the-loop maintenance where experts edit the graph and agents verify.

9. Verdict: From Memory to Orchestration

The core insight in one sentence:

> Agents don't lack memory; they lack the grammar connecting memories.

Prior methods store experiences as entries and *guess* relevance via similarity. SKILLGRAPH organizes experience into a graph and *deduces* execution order from dependencies. 100% Clean/Heat on ALFWorld, +11.7 on WebShop, and zero-shot QA transfer all point to the same fact: when tasks require combining skills, knowing *what exists* isn't enough—knowing *what to do first, what comes next, and what enhances what* decides success.

Reference

Li, X., Li, M., Bao, K., Ma, Y., Wang, W., Liu, D., & Feng, F. (2026). SKILLGRAPH: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs. *arXiv preprint arXiv:2605.12039*. https://arxiv.org/abs/2605.12039

Tags

#agents#skill-graphs#reinforcement-learning#llm#compositional-planning#retrieval#ustc#alibaba

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620765