English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GraSP Explained: When Skills Are No Longer the Bottleneck, Orchestration Is

Forum topic · 小凯 · 2026-04-26

Summary

This Chinese tech-forum post offers an in-depth analysis of GraSP (Graph-Structured Skill Compositions for LLM Agents), a paper from Tencent (arXiv 2604.17870v1). It begins with a counterintuitive finding: giving skill-based agents more skills can reduce performance, indicating the bottleneck has shifted from skill availability to skill orchestration. GraSP addresses this by compiling retrieved skills into a typed DAG with state, data, and order edges, then executing it topologically with precondition checks, postcondition verification, and five typed local-repair operators (REBIND, INSERTPREREQ, SUBSTITUTE, REWIRE, BYPASS). Repair is bounded to an h-hop neighborhood (O(d^h) instead of global replanning O(N)), with confidence-based routing falling back to ReAct as a no-regression guarantee. Across 48 (model, split) configurations on ALFWorld, ScienceWorld, WebShop, and InterCode using eight backbones (GPT-4.1, Claude, Gemini 2.5 Pro, DeepSeek V3.2, etc.), GraSP ranked first in all cells: +12.7 points over ExpeL and up to 41% fewer environment steps, with advantages growing with task complexity. The post also covers limitations, comparisons to ReAct, Reflexion, ExpeL, Voyager, and Synapse, and implications for agent frameworks like browser-use.

This is an English summary of a detailed Chinese forum post analyzing the paper GraSP: Graph-Structured Skill Compositions for LLM Agents (Tianle Xia, Lingxiang Hu, Yiding Sun et al., Tencent; arXiv 2604.17870v1 [cs.CL], 20 Apr 2026). Related project: https://github.com/browser-use/browser-harness (GraSP code pending open-source release).

Key points

The problem: less is more

  • Counterintuitive experimental result: more skills hurt agent performance. 2–3 focused skills yield the largest gains; 4+ skills show diminishing returns; comprehensive skill documentation can be actively harmful.
  • Two flaws in current skill-based agents:
  • 1. Cognitive overload — all retrieved skills are dumped into the prompt, forcing implicit in-context reasoning about which skill, in what order, under what conditions. 2. Lost causal structure — flat linear execution discards preconditions, effects, and dependencies. A failure at step k forces a full replan from scratch, O(N) complexity.
  • Root cause: a missing compilation layer between retrieval ("what is relevant") and execution ("run this step") that answers "how do these skills depend on each other?"
  • Architecture: a four-stage pipeline

    GraSP is a typed DAG G = (V, E) with nodes {v_src} ∪ V_skill ∪ {v_snk} and three edge types:

  • State edges — u's effect satisfies v's precondition (hard constraint)
  • Data edges — u's output binds to v's input (hard constraint)
  • Order edges — soft priorities / resource-conflict constraints
  • 1. Memory-conditioned retrieval: blends direct semantic matching with episodic memory from successful trajectories: p(s|q,x,R) = λ·p_dir(s|q,x) + (1-λ)·(1/Z)Σ ρ_j · freq(s, τ_ij). Retrieval confidence is calibrated from mean memory similarity, distribution consistency (1−JSD), top-skill margin, and goal coverage. 2. DAG compilation (the most critical stage): LLM proposes skill instances → parameter binding validation → edge inference from precondition-effect matching and memory priors → cycle resolution by dropping low-confidence soft edges → verifier attachment. Each node carries schema, parameters, pre/postconditions, verifier, state, confidence, and repair budget. 3. Verified execution with local repair: topological traversal with precondition checks, execution, and postcondition verification. Failures generate typed failure events handled by five repair operators: REBIND (fix parameters), INSERTPREREQ (insert missing-precondition subgraph), SUBSTITUTE (swap skill schema), REWIRE (edit edges), BYPASS (skip redundant node). Repair is bounded to an h-hop neighborhood; complexity O(d^h) (typically ~4–5 nodes) instead of O(N). 4. Confidence-based routing: high confidence → full DAG + local repair; medium → enlarged repair budget; low → ReAct fallback. Since a flat sequence is a DAG with only order edges, GraSP includes ReAct as a special case — a no-regression guarantee.

    Results: first place in all 48 configurations

  • Benchmarks: ALFWorld, ScienceWorld, WebShop, InterCode; 8 backbones (DeepSeek V3.2, GPT-4.1, Claude-4-Sonnet, GLM-5, Gemini 2.5 Pro, o4 Mini, Qwen3-235B, Kimi-K2.5); baselines: ReAct, Reflexion, ExpeL, ReAct+Skills.
  • 48/48 (model, split) cells ranked first: +12.7 points over ExpeL, +6.9 over the strongest per-cell baseline, up to +19 points, and ~24% fewer environment steps (max 41% on ScienceWorld unseen long-horizon tasks).
  • Ablations show every component contributes, with DAG compilation the most important; replacing local repair with global replanning loses 3.2/3.1 points.
  • Advantage grows with complexity: ~6% on short tasks (≤10 steps) vs ~18% on long tasks (≥20 steps).
  • Typed repair beats global replanning: precondition-failure recovery 84.2% vs 61.8% (+22.4%); postcondition-failure lead ~16%.
  • Robust to skill count and quality: flat execution peaks at M=3 and drops ~6% at M=8, while GraSP at M=8 (79.4) still beats flat's M=3 optimum (74.9). Degraded skill quality costs GraSP ~5% vs flat ~9%. DAG compilation acts as a structural filter that automatically excludes skills that cannot be connected via precondition-effect edges.
  • Why GraSP works

  • Skill execution is framed as a program compilation problem, not information retrieval: retrieval → "what is relevant"; compilation → "how dependencies are structured"; execution → "run now".
  • Exploits failure locality: a DAG failure only affects topological descendants, analogous to incremental compilation, transaction rollback, and checkpoint recovery.
  • Typed edges carry execution semantics (executability, parameter provenance, ordering), so repairs are structure-aware: knowing *why* something failed tells you how to fix it.
  • "Less is more" is explained by structure outperforming quantity: irrelevant skills distract flat agents, while DAG compilation selects a precise, connectable subset.
  • Limitations and open questions

  • Compilation itself costs LLM calls (possibly a negative optimization on simple tasks); dependence on accurate precondition/effect annotations; cycle resolution may drop important ordering; verifier design and human-in-the-loop scenarios under-discussed.
  • Versus related work: ReAct (no skill structure), Reflexion (reflects but doesn't restructure), ExpeL (learns but doesn't compile to graphs), Voyager (skill creation but flat execution), Synapse (skill graphs but no typed repair emphasis). GraSP's novelty is the complete compile → execute → repair loop.
  • Future directions: dynamic skill discovery during execution, probabilistic DAGs, cross-episode graph reuse as program templates, multi-agent graph merging, and neural compilers to cut compilation cost.

Takeaway

The next agent bottleneck is not more skills but better orchestration. GraSP upgrades skill execution from information retrieval to program compilation — typed DAGs plus local repair resolve the "more skills, worse performance" paradox, and structured orchestration beats a bigger skill library.

Tags

#llm-agents#grasp#skill-orchestration#dag#local-repair#paper-analysis#benchmark-results#agent-architecture

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618764