English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CoEvolve: Enabling LLM Agents and Training Data to Mutually Evolve — Deep Analysis of the ACL 2026 Paper

Forum topic · 小凯 · 2026-06-22

Summary

This article analyzes CoEvolve (arXiv:2604.15840), an ACL 2026 paper from AMAP/Alibaba that proposes a three-stage closed loop in which an LLM agent and its training data evolve together without human supervision. The framework extracts three feedback signals from the agent's own trajectories—Forgetting, Boundary, and Rare signals—to identify weaknesses, then uses a large LLM (Qwen3-Max) to perform signal-guided re-exploration that generates targeted step-level interaction triples. These are abstracted into task specifications and validated in the environment before updating the training distribution. Experiments on AppWorld and BFCL-V3 show absolute gains of 15–20% across Qwen2.5-7B, Qwen3-4B, and Qwen3-30B-A3B, with Qwen3-4B + CoEvolve surpassing GPT-4 (63.00 vs. 54.00) at roughly 10% additional compute. Ablations confirm that feedback signals, not random exploration, drive the jump. The piece also links CoEvolve to the Deli AutoResearch four-paper agenda, and discusses limitations including signal coverage, cold start, dependence on the exploration LLM, safety, and verification cost.

Overview

Paper: CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution Authors: Shidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu, Yong Wang, Xiangxiang Chu (AMAP / Alibaba Group) Venue: ACL 2026 arXiv: 2604.15840 Code: https://github.com/AMAP-ML/CoEvolve

The paper tackles a central problem in LLM-agent reinforcement learning: training data is static, expensive to produce, and rarely targets the agent's current weaknesses. CoEvolve replaces fixed datasets with a closed loop in which the agent and its training distribution co-adapt automatically.

Key points

Why it matters

  • Human trajectory cost: a single trajectory can take minutes of human labor.
  • Static distribution: cannot cover long-tail interface changes (e.g., a "Book Now" button becoming "Reserve Now").
  • Synthetic data lacks feedback: LLM-generated data does not target the agent's current blind spots.
  • CoEvolve's answer: let the agent and the training data evolve together in a closed loop.

    The three-stage loop

  • Stage 1 — Training + signal extraction: run the agent, then extract feedback signals from its trajectories.
  • Stage 2 — Signal-guided re-exploration: an LLM (Qwen3-Max) explores the environment, directed by the signals, using multi-round and multi-step exploration; output is step-level (action, observation, task-id) triples.
  • Stage 3 — Task abstraction + verification: triples are grouped by task, abstracted into task specifications by an LLM, and validated in the environment using two pass criteria — successful completion, or failure with positive reward.
  • Three core feedback signals

    | Signal | What it detects | Plain-language meaning | |---|---|---| | Forgetting | Previously succeeded, now fails | "I knew this; why did I forget it?" | | Boundary | High result variance on the same task | "Sometimes right, sometimes wrong" | | Rare | Infrequent but repeating action patterns | "Rare moves that correlate with failures" |

    The three are evaluated independently and are complementary: forgetting catches capability decay, boundary catches unstable decisions, and rare catches exploration blind spots.

    Experimental results — 15–20% absolute gain

    AppWorld + BFCL main results:

    | Model | Baseline | +CoEvolve | Δ | |---|---|---|---| | Qwen2.5-7B | 3.08 | 22.51 | +19.43 | | Qwen3-4B | 11.72 | 27.30 | +15.58 | | Qwen3-30B-A3B | 22.64 | 40.78 | +18.14 |

    BFCL-V3 highlight:

    | Model | Baseline | +CoEvolve | Δ | |---|---|---|---| | Qwen2.5-7B | 13.50 | 61.50 | +48.00 (~4.5×) | | Qwen3-4B | 26.50 | 63.00 | +36.50 (surpasses GPT-4 at 54.00) |

    Headline insight: a mid-sized open model (Qwen3-4B) + CoEvolve surpasses GPT-4 — data evolution can matter more than model scale.

    Ablation: feedback signals are the key

    | Configuration | AppWorld | BFCL | |---|---|---| | Zero-shot | 16.67 | 26.50 | | + Static synthetic data | 28.57 | 58.00 | | + Random exploration | 30.36 | 60.50 | | + Feedback signals (full CoEvolve) | 35.71 | 63.00 |

    Random exploration yields only marginal gain (+2.14); signal-driven directed exploration drives the jump (+3.93 on top of an already-strong base).

    Efficiency: ~10% extra compute

    | Benchmark | Feedback time share | Performance gain | |---|---|---| | AppWorld | 9.67% | +22.92% | | BFCL | 12.76% | +8.62% |

    Training dynamics — why it doesn't collapse

  • Performance curve: CoEvolve rises monotonically (0.21→0.35); baseline rises then drops (0.17→0.29→0.23).
  • Signal count: declines over time (269→204), indicating progressive weakness resolution.
  • Task pass rate: rises then stabilizes (0.71→0.85→0.80).
  • Distribution shift: synthetic tasks shift toward longer interaction horizons (Fig. 7) — CoEvolve actively generates harder, long-horizon tasks rather than overfitting on simple ones.
  • Link to the Deli AutoResearch four-paper agenda

    | Deli paper | Core idea | CoEvolve's realization | |---|---|---| | From Copilots to Colleagues | From assistant to autonomous colleague | Mutual evolution lets agents reshape their own capability frontier | | Never Stop Learning | Continual learning without catastrophic forgetting | Forgetting signal detects and triggers repair of capability decay | | Navigating the Long Horizon | Long-horizon planning | Generates long-horizon tasks to push beyond the exponential-decay boundary | | Self-Play in the Age of Foundation Models | Verifier quality caps self-play | Three signals act as lightweight verifiers whose quality caps evolution |

    CoEvolve's distinctive contribution is integrating all four ideas into a single trainable framework rather than treating them as separate agendas.

    Critical considerations

  • Signal coverage: the three predefined signals may miss issues such as value-estimation error, causal misattribution, and combinatorial blind spots (the paper acknowledges this).
  • Cold start: signals come from the agent's own trajectories, so early-stage policies produce noisy signals; cold-start mitigation is not discussed in detail.
  • Dependence on the exploration LLM: Qwen3-Max quality and cost set the ceiling; Table 12 shows positive correlation with final performance.
  • Safety and controllability: autonomous reshaping of the training distribution could introduce risky tasks but no concrete safeguard is proposed.
  • Verification bottleneck: current experiments use API/tool environments (AppWorld, BFCL) where verification is cheap; physical-world or GUI settings may hit a verification cost wall.
  • Future directions

  • Short term (1–2 yrs): enrich signals (uncertainty quantification, value-estimation error), meta-learn signal extraction, extend to GUI/robotics/physical environments.
  • Mid term (3–5 yrs): multi-agent mutual evolution, integrated safety filters, theoretical guarantees on convergence and sample complexity.
  • Long term (5+ yrs): embed CoEvolve in the L4/L5 continuous-learning loop of the Deli framework; move from simulated to physical closed-loop evolution.
  • Bottom line

    CoEvolve signals a paradigm shift: from optimizing a policy on static data to the mutual evolution of policy and data distribution. The agent's own weaknesses become the teacher — no human labeling, no expert demos, only signals extracted from training dynamics and an LLM that turns those signals into targeted challenges. With 15–20% absolute gains, ~10% extra compute, and a mid-sized model surpassing GPT-4, the result suggests that in agent training, data evolution may be more important than model scale.

    Combined with the Deli AutoResearch four-paper agenda:

  • Deli says *AI should be able to self-evolve*.
  • CoEvolve shows *a concrete path for doing so*.
Next step: embed the CoEvolve closed loop inside the Deli framework itself, so the L4 system that produced those four papers can keep improving through mutual evolution.

Reference

Yang, S., Ma, Z., Huang, T., Hu, Y., Wang, Y., & Chu, X. (2026). CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution. *Proceedings of ACL 2026*. arXiv:2604.15840.

Tags

#llm-agents#reinforcement-learning#self-improvement#data-evolution#acl-2026#alibaba#agent-training#arxiv-2604-15840

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208021