This is an English summary of a detailed Chinese forum post analyzing the paper GraSP: Graph-Structured Skill Compositions for LLM Agents (Tianle Xia, Lingxiang Hu, Yiding Sun et al., Tencent; arXiv 2604.17870v1 [cs.CL], 20 Apr 2026). Related project: https://github.com/browser-use/browser-harness (GraSP code pending open-source release).
Key points
The problem: less is more
- Counterintuitive experimental result: more skills hurt agent performance. 2–3 focused skills yield the largest gains; 4+ skills show diminishing returns; comprehensive skill documentation can be actively harmful.
- Two flaws in current skill-based agents: 1. Cognitive overload — all retrieved skills are dumped into the prompt, forcing implicit in-context reasoning about which skill, in what order, under what conditions. 2. Lost causal structure — flat linear execution discards preconditions, effects, and dependencies. A failure at step k forces a full replan from scratch, O(N) complexity.
- Root cause: a missing compilation layer between retrieval ("what is relevant") and execution ("run this step") that answers "how do these skills depend on each other?"
- State edges — u's effect satisfies v's precondition (hard constraint)
- Data edges — u's output binds to v's input (hard constraint)
- Order edges — soft priorities / resource-conflict constraints
- Benchmarks: ALFWorld, ScienceWorld, WebShop, InterCode; 8 backbones (DeepSeek V3.2, GPT-4.1, Claude-4-Sonnet, GLM-5, Gemini 2.5 Pro, o4 Mini, Qwen3-235B, Kimi-K2.5); baselines: ReAct, Reflexion, ExpeL, ReAct+Skills.
- 48/48 (model, split) cells ranked first: +12.7 points over ExpeL, +6.9 over the strongest per-cell baseline, up to +19 points, and ~24% fewer environment steps (max 41% on ScienceWorld unseen long-horizon tasks).
- Ablations show every component contributes, with DAG compilation the most important; replacing local repair with global replanning loses 3.2/3.1 points.
- Advantage grows with complexity: ~6% on short tasks (≤10 steps) vs ~18% on long tasks (≥20 steps).
- Typed repair beats global replanning: precondition-failure recovery 84.2% vs 61.8% (+22.4%); postcondition-failure lead ~16%.
- Robust to skill count and quality: flat execution peaks at M=3 and drops ~6% at M=8, while GraSP at M=8 (79.4) still beats flat's M=3 optimum (74.9). Degraded skill quality costs GraSP ~5% vs flat ~9%. DAG compilation acts as a structural filter that automatically excludes skills that cannot be connected via precondition-effect edges.
- Skill execution is framed as a program compilation problem, not information retrieval: retrieval → "what is relevant"; compilation → "how dependencies are structured"; execution → "run now".
- Exploits failure locality: a DAG failure only affects topological descendants, analogous to incremental compilation, transaction rollback, and checkpoint recovery.
- Typed edges carry execution semantics (executability, parameter provenance, ordering), so repairs are structure-aware: knowing *why* something failed tells you how to fix it.
- "Less is more" is explained by structure outperforming quantity: irrelevant skills distract flat agents, while DAG compilation selects a precise, connectable subset.
- Compilation itself costs LLM calls (possibly a negative optimization on simple tasks); dependence on accurate precondition/effect annotations; cycle resolution may drop important ordering; verifier design and human-in-the-loop scenarios under-discussed.
- Versus related work: ReAct (no skill structure), Reflexion (reflects but doesn't restructure), ExpeL (learns but doesn't compile to graphs), Voyager (skill creation but flat execution), Synapse (skill graphs but no typed repair emphasis). GraSP's novelty is the complete compile → execute → repair loop.
- Future directions: dynamic skill discovery during execution, probabilistic DAGs, cross-episode graph reuse as program templates, multi-agent graph merging, and neural compilers to cut compilation cost.
Architecture: a four-stage pipeline
GraSP is a typed DAG G = (V, E) with nodes {v_src} ∪ V_skill ∪ {v_snk} and three edge types:
1. Memory-conditioned retrieval: blends direct semantic matching with episodic memory from successful trajectories: p(s|q,x,R) = λ·p_dir(s|q,x) + (1-λ)·(1/Z)Σ ρ_j · freq(s, τ_ij). Retrieval confidence is calibrated from mean memory similarity, distribution consistency (1−JSD), top-skill margin, and goal coverage.
2. DAG compilation (the most critical stage): LLM proposes skill instances → parameter binding validation → edge inference from precondition-effect matching and memory priors → cycle resolution by dropping low-confidence soft edges → verifier attachment. Each node carries schema, parameters, pre/postconditions, verifier, state, confidence, and repair budget.
3. Verified execution with local repair: topological traversal with precondition checks, execution, and postcondition verification. Failures generate typed failure events handled by five repair operators: REBIND (fix parameters), INSERTPREREQ (insert missing-precondition subgraph), SUBSTITUTE (swap skill schema), REWIRE (edit edges), BYPASS (skip redundant node). Repair is bounded to an h-hop neighborhood; complexity O(d^h) (typically ~4–5 nodes) instead of O(N).
4. Confidence-based routing: high confidence → full DAG + local repair; medium → enlarged repair budget; low → ReAct fallback. Since a flat sequence is a DAG with only order edges, GraSP includes ReAct as a special case — a no-regression guarantee.
Results: first place in all 48 configurations
Why GraSP works
Limitations and open questions
Takeaway
The next agent bottleneck is not more skills but better orchestration. GraSP upgrades skill execution from information retrieval to program compilation — typed DAGs plus local repair resolve the "more skills, worse performance" paradox, and structured orchestration beats a bigger skill library.