Skill-MAS: Evolving Meta-Skills for Automatic Multi-Agent Systems
TL;DR: Ant Group and HKUST(GZ) propose Skill-MAS — treating multi-agent orchestration strategy as an evolvable "Meta-Skill," letting frozen LLMs continuously learn without parameter updates via multi-trajectory sampling and selective reflection. A frozen GPT-5.4-Nano optimized with Skill-MAS outperforms expensive test-time re-ranking strategies, and the learned meta-skills transfer across models and tasks.
The "Impossible Triangle" of Multi-Agent Orchestration
| Approach | Model Capability | Experience Accumulation | Core Problem | |------|---------|---------|---------| | Test-time orchestration | Strong (GPT-5, Claude, etc.) | ❌ None | Starts from scratch each time, repeats mistakes | | Training-time orchestration | Weak (7B models only) | ✓ Yes | Cannot scale to frontier models; huge data needs |
Test-time orchestration means a frontier LLM orchestrates agents but never learns from experience. Training-time orchestration internalizes orchestration ability via gradient updates — infeasible for >100B models.
Skill-MAS asks: is there a third path — retaining frontier-model reasoning while accumulating experience like training-based methods?
Answer: externalize the orchestration strategy as an evolvable natural-language artifact, executed by a frozen LLM.
Meta-Skill: Strategy as Code
Skill-MAS's core innovation is the Meta-Skill — like a director (frontier LLM) who writes a directing manual, consults it before each shoot, and updates it after. The director (model parameters) never changes; the manual (meta-skill) keeps evolving.
The Meta-Skill is a structured natural-language document with three modules:
Module 1: Task Decomposition (What)
- Intent and scope analysis
- Decomposition into logically coherent subtasks
- Dependency mapping
- Success criteria definition
- Role profiles with distinct identities
- Precise instruction design
- Input context framing
- Topology selection (serial, hierarchical, loop, etc.)
- Data flow and state management
- Executable code generation
- Uncertainty \(u_i\) = standard deviation of scores
- Difficulty \(d_i\) = -mean score
- *Within-task*: contrast high-score vs low-score trajectories to find local patterns
- *Cross-task*: distill systematic patterns into a prioritized evidence package \(E\)
- Cross-model: a Meta-Skill evolved on GPT-5.4-Nano lifts Qwen3.5-Plus from 18.45 to 24.40 — learned strategies are not model-specific quirks.
- Cross-task: BCP-evolved skill lifts VitaBench from 0 to 13.10; VitaBench→BCP goes from 0 to 23.21.
- Cross-model + cross-task (hardest): still positive transfer.
- Selective reflection needs ground-truth labels; LLM-as-a-judge could enable self-supervised evaluation.
- Multi-task learning is not yet systematically optimized.
- Deeper question: if all orchestration strategy can be externalized as natural language, where does LLM "intelligence" reside? Skill-MAS's answer: both matter, but strategy can evolve independently of the executor — decoupling "algorithm" from "hardware."
- Lin et al. (2026). Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems. *Ant Group & HKUST(GZ)*, arXiv:2606.18837.
Module 2: Agent Engineering (Who)
Module 3: Workflow Orchestration (How)
This echoes OpenClaw's Skill mechanism — externalizing strategy as version-controlled text rather than burying it in model weights.
Closed-Loop Evolution: Multi-Trajectory Sampling + Selective Reflection
Stage 1: Multi-Trajectory Rollout
For each task under the current meta-skill \(S^(r)\), sample \(K=5\) independent execution trajectories. Multiple rollouts separate "true capability" from "execution noise," yielding:
Stage 2: Selective Reflection
1. Priority-driven task selection: fuse \(u_i\) and \(d_i\) into a priority score, find the "elbow" on the priority curve, and reflect only on the top \(j^*\) tasks — those both hard and uncertain.
2. Hierarchical trajectory reflection:
3. Skill optimization: targeted updates to the three modules (patch weaknesses, keep the scaffold, abstract into generalizable principles), producing \(S^(r+1)\).
Experiments: Four Models, Four Complex Domains
Benchmarks
| Benchmark | Domain | Metric | |------|------|---------| | DeepResearchBench (DRB) | Deep research reports | Comprehensiveness, insight, instruction-following, readability | | Humanity's Last Exam-Math (HLE) | Expert-level math | Accuracy | | BrowseComp-Plus (BCP) | Multi-hop dynamic QA | Accuracy | | VitaBench | Real-world interactive tool use | Rubric-based success rate |
Models: Gemini-3.1-Flash, GPT-5.4-Nano, Qwen3.5-Plus, DeepSeek-V4-Flash
Key Results
Skill-MAS achieved the highest average performance on all four models:
| Model | Best Baseline | Skill-MAS Initial | Skill-MAS Optimized | Gain | |------|--------------|-------------------|---------------------|------| | Gemini-3.1-Flash | 21.29 | 21.68 | 29.49 | +38.5% | | GPT-5.4-Nano | 24.83 | 19.64 | 27.55 | +11.0% | | Qwen3.5-Plus | 32.23 | 32.61 | 38.41 | +19.2% | | DeepSeek-V4-Flash | 35.70 | 33.72 | 41.05 | +15.0% |
Cost-performance trade-off: training-time MAS is cheapest but weakest; test-time MAS is strong but priciest; Skill-MAS is highest-performing at moderate cost (one-time meta-skill evolution, zero extra cost afterward).
Transferability: Learning Orchestration Wisdom, Not Domain Tricks
A Real Evolution Trajectory: From Chaos to Order
On BrowseComp-Plus, the meta-skill evolved: enriched decomposition with evidence weighting → agent-level epistemic control (weighted satisfaction protocols) → system-level resilience (backtracking, link verification, node re-execution). This is directional, layered strategy accumulation — not random search.
Implications for Agent Developers
1. Strategy externalization > parameter internalization: write orchestration strategies as version-controlled structured documents (SKILL.md style). 2. Multi-trajectory sampling is essential: K=5 statistics tell you whether the model "really knows it or got lucky." 3. Selective reflection > full retrospective: prioritize tasks by uncertainty × difficulty. 4. Cross-level abstraction: distill lessons into general principles — the key to cross-task transfer.