Overview
Paper: CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution Authors: Shidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu, Yong Wang, Xiangxiang Chu (AMAP / Alibaba Group) Venue: ACL 2026 arXiv: 2604.15840 Code: https://github.com/AMAP-ML/CoEvolve
The paper tackles a central problem in LLM-agent reinforcement learning: training data is static, expensive to produce, and rarely targets the agent's current weaknesses. CoEvolve replaces fixed datasets with a closed loop in which the agent and its training distribution co-adapt automatically.
Key points
Why it matters
- Human trajectory cost: a single trajectory can take minutes of human labor.
- Static distribution: cannot cover long-tail interface changes (e.g., a "Book Now" button becoming "Reserve Now").
- Synthetic data lacks feedback: LLM-generated data does not target the agent's current blind spots.
- Stage 1 — Training + signal extraction: run the agent, then extract feedback signals from its trajectories.
- Stage 2 — Signal-guided re-exploration: an LLM (Qwen3-Max) explores the environment, directed by the signals, using multi-round and multi-step exploration; output is step-level (action, observation, task-id) triples.
- Stage 3 — Task abstraction + verification: triples are grouped by task, abstracted into task specifications by an LLM, and validated in the environment using two pass criteria — successful completion, or failure with positive reward.
- Performance curve: CoEvolve rises monotonically (0.21→0.35); baseline rises then drops (0.17→0.29→0.23).
- Signal count: declines over time (269→204), indicating progressive weakness resolution.
- Task pass rate: rises then stabilizes (0.71→0.85→0.80).
- Distribution shift: synthetic tasks shift toward longer interaction horizons (Fig. 7) — CoEvolve actively generates harder, long-horizon tasks rather than overfitting on simple ones.
- Signal coverage: the three predefined signals may miss issues such as value-estimation error, causal misattribution, and combinatorial blind spots (the paper acknowledges this).
- Cold start: signals come from the agent's own trajectories, so early-stage policies produce noisy signals; cold-start mitigation is not discussed in detail.
- Dependence on the exploration LLM: Qwen3-Max quality and cost set the ceiling; Table 12 shows positive correlation with final performance.
- Safety and controllability: autonomous reshaping of the training distribution could introduce risky tasks but no concrete safeguard is proposed.
- Verification bottleneck: current experiments use API/tool environments (AppWorld, BFCL) where verification is cheap; physical-world or GUI settings may hit a verification cost wall.
- Short term (1–2 yrs): enrich signals (uncertainty quantification, value-estimation error), meta-learn signal extraction, extend to GUI/robotics/physical environments.
- Mid term (3–5 yrs): multi-agent mutual evolution, integrated safety filters, theoretical guarantees on convergence and sample complexity.
- Long term (5+ yrs): embed CoEvolve in the L4/L5 continuous-learning loop of the Deli framework; move from simulated to physical closed-loop evolution.
- Deli says *AI should be able to self-evolve*.
- CoEvolve shows *a concrete path for doing so*.
CoEvolve's answer: let the agent and the training data evolve together in a closed loop.
The three-stage loop
Three core feedback signals
| Signal | What it detects | Plain-language meaning | |---|---|---| | Forgetting | Previously succeeded, now fails | "I knew this; why did I forget it?" | | Boundary | High result variance on the same task | "Sometimes right, sometimes wrong" | | Rare | Infrequent but repeating action patterns | "Rare moves that correlate with failures" |
The three are evaluated independently and are complementary: forgetting catches capability decay, boundary catches unstable decisions, and rare catches exploration blind spots.
Experimental results — 15–20% absolute gain
AppWorld + BFCL main results:
| Model | Baseline | +CoEvolve | Δ | |---|---|---|---| | Qwen2.5-7B | 3.08 | 22.51 | +19.43 | | Qwen3-4B | 11.72 | 27.30 | +15.58 | | Qwen3-30B-A3B | 22.64 | 40.78 | +18.14 |
BFCL-V3 highlight:
| Model | Baseline | +CoEvolve | Δ | |---|---|---|---| | Qwen2.5-7B | 13.50 | 61.50 | +48.00 (~4.5×) | | Qwen3-4B | 26.50 | 63.00 | +36.50 (surpasses GPT-4 at 54.00) |
Headline insight: a mid-sized open model (Qwen3-4B) + CoEvolve surpasses GPT-4 — data evolution can matter more than model scale.
Ablation: feedback signals are the key
| Configuration | AppWorld | BFCL | |---|---|---| | Zero-shot | 16.67 | 26.50 | | + Static synthetic data | 28.57 | 58.00 | | + Random exploration | 30.36 | 60.50 | | + Feedback signals (full CoEvolve) | 35.71 | 63.00 |
Random exploration yields only marginal gain (+2.14); signal-driven directed exploration drives the jump (+3.93 on top of an already-strong base).
Efficiency: ~10% extra compute
| Benchmark | Feedback time share | Performance gain | |---|---|---| | AppWorld | 9.67% | +22.92% | | BFCL | 12.76% | +8.62% |
Training dynamics — why it doesn't collapse
Link to the Deli AutoResearch four-paper agenda
| Deli paper | Core idea | CoEvolve's realization | |---|---|---| | From Copilots to Colleagues | From assistant to autonomous colleague | Mutual evolution lets agents reshape their own capability frontier | | Never Stop Learning | Continual learning without catastrophic forgetting | Forgetting signal detects and triggers repair of capability decay | | Navigating the Long Horizon | Long-horizon planning | Generates long-horizon tasks to push beyond the exponential-decay boundary | | Self-Play in the Age of Foundation Models | Verifier quality caps self-play | Three signals act as lightweight verifiers whose quality caps evolution |
CoEvolve's distinctive contribution is integrating all four ideas into a single trainable framework rather than treating them as separate agendas.
Critical considerations
Future directions
Bottom line
CoEvolve signals a paradigm shift: from optimizing a policy on static data to the mutual evolution of policy and data distribution. The agent's own weaknesses become the teacher — no human labeling, no expert demos, only signals extracted from training dynamics and an LLM that turns those signals into targeted challenges. With 15–20% absolute gains, ~10% extra compute, and a mid-sized model surpassing GPT-4, the result suggests that in agent training, data evolution may be more important than model scale.
Combined with the Deli AutoResearch four-paper agenda:
Reference
Yang, S., Ma, Z., Huang, T., Hu, Y., Wang, Y., & Chu, X. (2026). CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution. *Proceedings of ACL 2026*. arXiv:2604.15840.