English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ctx2Skill: LLMs Evolve Reusable Skills from Context via Multi-Agent Self-Play

Forum topic · 小凯 · 2026-05-16

Summary

Ctx2Skill is a framework that lets large language models autonomously extract reusable skills from complex, unseen contexts through a multi-agent self-play loop. A Challenger generates probing tasks with rubrics, a Reasoner attempts them, and a fixed Judge returns binary pass/fail feedback. Proposer-Generator pairs then diagnose failures to update the Reasoner's skills and analyze successes to make the Challenger harder, creating continuous co-evolutionary pressure. A Cross-Time Replay mechanism selects the most balanced skill set across all iterations using hard and easy probes, guarding against adversarial collapse. On CL-bench (500 contexts, 1,899 tasks, 31,607 rubrics), Ctx2Skill raises GPT-4.1 from 11.1% to 16.5% and GPT-5.1 from 21.1% to 25.8%, outperforming prompting and windowed extraction baselines. Skills are human-readable, editable, and transfer across models—strong-model skills transfer well to weaker models, though not vice versa. The framework offers a scalable, annotation-free paradigm for in-context learning.

Overview

Ctx2Skill (arXiv:2604.27660; Tsinghua University, DeepLang AI, UIUC, Fudan, CUHK; project: https://github.com/S1s-Z/Ctx2Skill) addresses a key weakness of current LLMs: in-context learning on novel, complex documents (technical manuals, experimental data, domain papers) that go beyond pretraining knowledge.

Why existing approaches fail

  • Manual annotation of long documents into skills is prohibitively expensive.
  • Automatic extraction lacks feedback: unlike code or math, there is no ground truth to verify whether extracted skills are complete or faithful.
  • The Multi-Agent Self-Play Loop

    Five roles iterate:

  • Challenger — generates probing tasks + grading rubrics from the context; evolves to stay adversarial
  • Reasoner — attempts tasks using its current skill set; evolves to fix knowledge gaps
  • Judge — fixed GPT-5.1, binary pass/fail verdicts
  • Proposer (one per side) — diagnoses failure cases (feeds Reasoner updates) or success cases (feeds Challenger updates)
  • Generator (one per side) — converts diagnoses into concrete skill additions/deletions/merges
  • Critically, the Challenger also evolves; without it, tasks become trivial and the Reasoner stops exposing blind spots. Updates are failure-driven text edits, and neither side sees the other's skills.

    Cross-Time Replay: Preventing Adversarial Collapse

    Self-play risks the Challenger drifting to degenerate tasks and the Reasoner over-specializing. Cross-Time Replay instead selects the best skill set from all historical iterations using training-free probes:

  • Hard probe: the failed task passing the fewest rubrics each round
  • Easy probe: the successful task passing the fewest rubrics each round
  • Selection: argmax_i (ρ^h(i) · ρ^e(i)) with Laplace smoothing — the product penalizes skill sets that solve hard tasks by sacrificing easy ones (or vice versa)
  • Empirically, fixed-iteration skill quality degrades monotonically (Iter-1: 15.9% → Iter-5: 14.7%), while Cross-Time Replay achieves 16.5%.

    Results on CL-bench

    CL-bench: 500 complex contexts, 1,899 tasks, 31,607 validation rubrics; four categories (domain knowledge reasoning, rule system application, procedural execution, empirical discovery & simulation); average 10.4K tokens (max 65K); 51.1% multi-turn sequential tasks.

    | Model | Baseline | Ctx2Skill | Gain | |---|---|---|---| | GPT-4.1 | 11.1% | 16.5% | +5.4% | | GPT-5.1 | 21.1% | 25.8% | +4.6% | | GPT-5.2 | 18.2% | 21.4% | +3.2% |

  • Baselines: single-shot skill prompting (+1.2% on GPT-4.1), windowed extraction (AutoSkill4Doc-style, +2.1%) — both far below Ctx2Skill.
  • GPT-4.1 + Ctx2Skill (16.5%) surpasses un-augmented Gemini 3 Pro (15.8%) — skills can bridge model capability gaps.
  • GPT-4.1-judged quality: Ctx2Skill skills score highest on conciseness (85.2), faithfulness (84.8), clarity (96.2), effectiveness (90.5), and reusability (92.5).
  • Ablations (GPT-4.1 overall, 16.5% full)

  • w/o Challenger evolution: 13.8% (largest drop — adversarial pressure is essential)
  • w/o Cross-Time Replay: 14.7%
  • w/o Hard Probe: 15.2%; w/o Easy Probe: 15.7%; w/o Laplace smoothing: 15.5%
  • Merging Proposer+Generator: 15.9%
  • Cross-Model Skill Transfer

  • GPT-5.1 skills → GPT-4.1: 16.1% (≈ self-produced 16.5%)
  • GPT-4.1 skills → GPT-5.1: 23.1% (< self-produced 25.8%)
  • This asymmetry suggests skill discovery itself may be an emergent capability of stronger models.

    Analysis Highlights

  • Skills, not memory: natural-language skill files are interpretable, editable, model-agnostic, and require no parameter access — friendly to closed-source APIs.
  • Minimal viable feedback: binary pass/fail plus adversarial pressure forms a weak but effective learning signal without ground truth.
  • Diminishing returns for stronger models: absolute gains shrink from +5.4% (GPT-4.1) to +3.2% (GPT-5.2), implying the bottleneck shifts toward precise knowledge extraction.
  • Limitations

  • High compute cost (5 iterations × 5 tasks × multiple agents per context)
  • Feedback quality depends on the fixed Judge (GPT-5.1)
  • Assumes all needed knowledge is in the given context (no external retrieval)
  • Static skills don't adapt dynamically across sequential task dependencies
  • Cross-Time Replay mitigates but does not eliminate adversarial collapse
  • Applications

    Enterprise knowledge bases (auto-extracting operating procedures), research assistants (inducing methodology from papers/data), and education (extracting problem-solving strategies from textbooks).

    References

  • Si, S. et al. (2026). *From Context to Skills: Can Language Models Learn from Context Skillfully?* arXiv:2604.27660.
  • Dou, S. et al. (2026). CL-bench. arXiv:2602.03587.
  • Yang, Y. et al. (2026). AutoSkill. arXiv:2603.01145.
  • Zhang, H. et al. (2026). CoEvoSkills. arXiv:2604.01687.

Tags

#ctx2skill#in-context-learning#multi-agent#self-play#skill-extraction#llm#benchmark#cross-time-replay

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620123