English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CoEvoSkills: AI Self-Evolving Agent Skills via Three-Party Co-Evolutionary Verification

Forum topic · 小凯 · 2026-05-30

Summary

CoEvoSkills (Self-Evolving Agent Skills via Co-Evolutionary Verification) is an April 2026 arXiv paper from researchers at UIC, MBZUAI, McGill, Columbia, Zhejiang, and UBC. It addresses a key problem in Anthropic's Agent Skills paradigm: human-written skill packages show inconsistent or even negative gains for LLM agents due to human-machine cognitive misalignment. The paper proposes a three-party co-evolutionary framework with information isolation: (1) a Skill Generator that iteratively produces multi-file skill packages, (2) a Surrogate Verifier running in a fully separate LLM session that creates test assertions and structured failure diagnostics without seeing generator internals, and (3) a Ground Truth Oracle that re-executes skills in fresh environments and returns only binary pass/fail signals to prevent test overfitting. On SkillsBench (87 tasks, 11 domains), CoEvoSkills achieves 71.1% pass rate versus 30.6% for the no-skill baseline (+40.5pp), surpassing human-written skills within 5 evolution rounds. Removing the Verifier drops performance by 30pp. Evolved skill packages transfer across six different LLM vendors with +35-44pp improvements, demonstrating that AI-generated skills capture reasoning patterns LLMs actually need.

CoEvoSkills: Deep Research on Three-Party Co-Evolution for Self-Evolving Agent Skills

CoEvoSkills (Self-Evolving Agent Skills via Co-Evolutionary Verification) is a paper published on arXiv in April 2026, with authors from six universities including UIC, MBZUAI, McGill, Columbia, Zhejiang, and UBC.

Core problem: After Anthropic introduced the Agent Skills concept, manually writing skills proved labor-intensive and suffers from human-machine cognitive misalignment — guides written by human experts sometimes work *worse* when AI uses them. Can AI evolve its own skills?

The paper's answer: through a three-party game co-evolutionary framework, AI can autonomously generate skill packages better than human-written ones.

From Tool to Skill: Why a Single Function Isn't Enough

Existing LLM Agent Tools are single function calls — check weather, do math, read files. But real-world professional tasks are multi-step, multi-file, cross-tool orchestration problems: fixing complex software bugs requires reading code, writing tests, running debuggers, verifying fixes; scientific analysis requires loading data, cleaning, modeling, visualizing, writing reports. These tasks need a structured workflow, not a single tool call.

Anthropic's Agent Skills were designed for this: a Skill isn't a single function but a multi-file structured package containing:

  • workflow instructions
  • executable scripts
  • domain references
  • This is like a micro project template — the Agent receives the package and knows how to complete a class of tasks.

    Why Human-Written Skills Sometimes Make AI Worse

    The paper tested carefully human-crafted skills on SkillsBench (87 tasks, 11 professional domains). Results were counterintuitive:

  • Some domains (e.g., software engineering) showed clear gains
  • Natural Science domains showed negative returns — adding the skill package made AI perform worse than without it
  • The authors attribute this to human-machine cognitive misalignment: workflows and abstraction levels designed by human experts are for human understanding and execution, but LLM Agents' reasoning style, context handling, and tool-calling habits differ fundamentally from humans. Steps that feel "intuitive" to humans may be information overload or wrongly sequenced for AI.

    This raises a fundamental question: if human-written skills aren't necessarily good for AI, who is better suited to write them?

    The Three-Party Game: Student, Mentor, and Ruthless Examiner

    CoEvoSkills consists of three information-isolated components:

    ① Skill Generator (the student)

    Generates and iterates on skill packages. Starting from task instructions, it produces a candidate skill package, executes it, and gets outputs. It then reads the verifier's feedback and improves the skill in the next iteration. Crucially, it maintains a persistent conversation context accumulating all historical verification feedback — each new skill version builds on prior failure lessons.

    ② Surrogate Verifier (the mentor)

    Runs in a completely separate LLM session and cannot see the generator's code, reasoning, or skill content. It only sees task instructions and the generator's output files. It generates test assertions and scores the outputs. On test failure, it produces structured failure diagnostics (which assertions failed, root cause analysis, actionable fix suggestions) fed back to the generator.

    Information isolation is the soul of this design. If the verifier could see the generator's code, it would inherit the generator's biases — what the generator thinks is right, the verifier would also think is right. This is confirmation bias, the fatal flaw of self-verification systems. Information isolation ensures the verifier provides independent, external, potentially contradictory feedback.

    ③ Ground Truth Oracle (the ruthless examiner)

    Re-executes the skill package independently in a fresh environment and returns only a binary pass/fail signal. No test contents, failure details, or scoring criteria are revealed. This prevents the generator from overfitting to the tests — if it knew what the tests check, it might optimize only for "passing tests" rather than "solving the problem correctly."

    Alternating Optimization: Two Feedback Loops Drive Evolution

    The process is alternating optimization:

    1. Generator produces skill → executes → outputs to Verifier 2. Verifier tests: if fail → diagnostic feedback → Generator improves skill (fixed test suite) 3. Verifier tests: if pass → hand off to Oracle for validation 4. If Oracle fails → returns only "fail" (no details) → Verifier must independently upgrade its test suite (generating stricter, more diverse tests) 5. Back to step 1, looping

    The elegance: Generator and Verifier pressure each other. The Generator grows stronger under the Verifier's testing pressure; the Verifier escalates test difficulty based on the Oracle's "escaped" signals. Both co-evolve, converging within 5 rounds.

    Core Results: The Numbers Speak

    | Condition | SkillsBench Pass Rate | |------|-------------------| | No-skill baseline | 30.6% | | Background knowledge only (no evolution) | 48.6% | | Verifier removed (Oracle's opaque signal only) | 41.1% | | CoEvoSkills (full framework) | 71.1% | | Carefully human-written skills | Mixed (negative returns in some domains) |

    Key findings:

  • +40.5pp over the no-skill baseline
  • Surpasses human-written skills within 5 evolution rounds
  • Removing the Verifier drops 30pp: diagnostic feedback is the fuel of evolution; without it, the generator gropes in the dark
  • Cross-model transfer: the same evolved skill package transferred to 6 LLMs from different vendors yields +35-44pp improvements
---

> Core insight: CoEvoSkills' key insight is not "letting AI write its own skills," but letting AI write skills in the AI way. Human-written skills fail not because humans aren't smart enough, but because human cognitive structures don't match LLMs. Skills that AI evolves itself capture the reasoning patterns and tool-usage strategies LLMs actually need, not the steps humans think "should" exist.

Tags

#ai-agents#llm#agent-skills#co-evolution#arxiv#automated-verification#skillsbench#research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980605