Overview
Ctx2Skill (arXiv:2604.27660; Tsinghua University, DeepLang AI, UIUC, Fudan, CUHK; project: https://github.com/S1s-Z/Ctx2Skill) addresses a key weakness of current LLMs: in-context learning on novel, complex documents (technical manuals, experimental data, domain papers) that go beyond pretraining knowledge.
Why existing approaches fail
- Manual annotation of long documents into skills is prohibitively expensive.
- Automatic extraction lacks feedback: unlike code or math, there is no ground truth to verify whether extracted skills are complete or faithful.
- Challenger — generates probing tasks + grading rubrics from the context; evolves to stay adversarial
- Reasoner — attempts tasks using its current skill set; evolves to fix knowledge gaps
- Judge — fixed GPT-5.1, binary pass/fail verdicts
- Proposer (one per side) — diagnoses failure cases (feeds Reasoner updates) or success cases (feeds Challenger updates)
- Generator (one per side) — converts diagnoses into concrete skill additions/deletions/merges
- Hard probe: the failed task passing the fewest rubrics each round
- Easy probe: the successful task passing the fewest rubrics each round
- Selection:
argmax_i (ρ^h(i) · ρ^e(i))with Laplace smoothing — the product penalizes skill sets that solve hard tasks by sacrificing easy ones (or vice versa) - Baselines: single-shot skill prompting (+1.2% on GPT-4.1), windowed extraction (AutoSkill4Doc-style, +2.1%) — both far below Ctx2Skill.
- GPT-4.1 + Ctx2Skill (16.5%) surpasses un-augmented Gemini 3 Pro (15.8%) — skills can bridge model capability gaps.
- GPT-4.1-judged quality: Ctx2Skill skills score highest on conciseness (85.2), faithfulness (84.8), clarity (96.2), effectiveness (90.5), and reusability (92.5).
- w/o Challenger evolution: 13.8% (largest drop — adversarial pressure is essential)
- w/o Cross-Time Replay: 14.7%
- w/o Hard Probe: 15.2%; w/o Easy Probe: 15.7%; w/o Laplace smoothing: 15.5%
- Merging Proposer+Generator: 15.9%
- GPT-5.1 skills → GPT-4.1: 16.1% (≈ self-produced 16.5%)
- GPT-4.1 skills → GPT-5.1: 23.1% (< self-produced 25.8%)
- Skills, not memory: natural-language skill files are interpretable, editable, model-agnostic, and require no parameter access — friendly to closed-source APIs.
- Minimal viable feedback: binary pass/fail plus adversarial pressure forms a weak but effective learning signal without ground truth.
- Diminishing returns for stronger models: absolute gains shrink from +5.4% (GPT-4.1) to +3.2% (GPT-5.2), implying the bottleneck shifts toward precise knowledge extraction.
- High compute cost (5 iterations × 5 tasks × multiple agents per context)
- Feedback quality depends on the fixed Judge (GPT-5.1)
- Assumes all needed knowledge is in the given context (no external retrieval)
- Static skills don't adapt dynamically across sequential task dependencies
- Cross-Time Replay mitigates but does not eliminate adversarial collapse
- Si, S. et al. (2026). *From Context to Skills: Can Language Models Learn from Context Skillfully?* arXiv:2604.27660.
- Dou, S. et al. (2026). CL-bench. arXiv:2602.03587.
- Yang, Y. et al. (2026). AutoSkill. arXiv:2603.01145.
- Zhang, H. et al. (2026). CoEvoSkills. arXiv:2604.01687.
The Multi-Agent Self-Play Loop
Five roles iterate:
Critically, the Challenger also evolves; without it, tasks become trivial and the Reasoner stops exposing blind spots. Updates are failure-driven text edits, and neither side sees the other's skills.
Cross-Time Replay: Preventing Adversarial Collapse
Self-play risks the Challenger drifting to degenerate tasks and the Reasoner over-specializing. Cross-Time Replay instead selects the best skill set from all historical iterations using training-free probes:
Empirically, fixed-iteration skill quality degrades monotonically (Iter-1: 15.9% → Iter-5: 14.7%), while Cross-Time Replay achieves 16.5%.
Results on CL-bench
CL-bench: 500 complex contexts, 1,899 tasks, 31,607 validation rubrics; four categories (domain knowledge reasoning, rule system application, procedural execution, empirical discovery & simulation); average 10.4K tokens (max 65K); 51.1% multi-turn sequential tasks.
| Model | Baseline | Ctx2Skill | Gain | |---|---|---|---| | GPT-4.1 | 11.1% | 16.5% | +5.4% | | GPT-5.1 | 21.1% | 25.8% | +4.6% | | GPT-5.2 | 18.2% | 21.4% | +3.2% |
Ablations (GPT-4.1 overall, 16.5% full)
Cross-Model Skill Transfer
This asymmetry suggests skill discovery itself may be an emergent capability of stronger models.
Analysis Highlights
Limitations
Applications
Enterprise knowledge bases (auto-extracting operating procedures), research assistants (inducing methodology from papers/data), and education (extracting problem-solving strategies from textbooks).