When AI Agents Learn to Team Up but Never Learn to Stop
On May 4, 2026, Chenchen Zhang posted a survey that is quietly disturbing. He systematically cataloged every academic paper on building multi-agent systems with LLMs and training them with reinforcement learning, tagging each by reward type, credit granularity, orchestration topology, and application domain — 84 papers in total, spanning 2022 through May 2026.
The finding:
> Of 84 papers, the number studying how multi-agent systems learn to stop is 0.
Not few. Zero.
Key points
- The paper proposes orchestration trace: a temporal event graph recording all coordination decisions (spawn, delegate, communicate, aggregate, stop) — as opposed to a single agent's chain-of-thought.
- Across three analytical axes, critical sparsity emerges:
- Rewards: 10 reward types identified; hybrid is most common (15 papers), but orchestration rewards appear in only 7 papers. Most work optimizes individual agent correctness, not system-level coordination:
- Credit assignment: 8 granularities from token to team. Agent-level dominates (23 papers); message-level credit appears in only 2 papers (2.4%) — nobody can answer which message caused which outcome, since the message-to-reward gradient is disconnected.
- Orchestration learning: decomposed into 5 atomic decisions — Spawn (O1), Delegate (O2), Communicate (O3), Aggregate (O4), Stop (O5). O5 has zero papers.
- Academia–industry scale gap: industrial systems (Moonshot AI's Kimi Agent Swarm, OpenAI Codex, Anthropic Claude Code) deploy large-scale orchestrated agents with dynamic stopping logic, but publish almost no training details. Per the paper: "The resulting scale gap is a gap between publicly reported deployment envelopes and open academic evaluation regimes, not independent verification of industrial training traces."
In 77 of 84 papers, this term is effectively zero.
Why stopping matters
Without a learned stopping policy, three failure modes occur:
1. Premature stopping — a fix is committed before tests are written. 2. Late stopping — endless "additional checks" waste compute and degrade user experience (overthinking / token black hole). 3. Never stopping — circular dependency loops among agents that only hard-coded round caps break.
A learned stop condition would require:
Today all of these Q functions are hard-coded heuristics, not trained.
Proposed directions
1. Explicit RL objective functions for O5, e.g., a value function \(V_{stop}(s_t)\) comparing continue-vs-stop expected returns. 2. Recording stop signals (and rationale) inside orchestration traces for future learning. 3. Cross academic–industrial evaluation standards for orchestration quality. 4. Message-level credit assignment methods requiring new gradient estimation over discrete message-reward paths.
Paper details
| Item | Value | |:-----|:------| | Title | Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces | | Author | Chenchen Zhang | | arXiv | 2605.02801 | | Date | May 4, 2026 | | Code | github.com/xxzcc/awesome-llm-mas-rl |
Corpus (84 retained papers): RL methods 42, benchmarks 18, classical MARL foundations 10, industrial reports 6, surveys 5, frameworks 3 (116 screened, 32 excluded).
Six orchestration topologies: centralized orchestrator + sub-agents; planner-executor-critic; debate/committee; parallel swarm; hierarchical; harness.
Bottom line
Multi-agent systems are moving from research toys to industrial infrastructure. An orchestra where every musician is trained but the conductor never learned to end the piece is a warning — 84 papers, zero on stopping, is a structural blind spot, not a statistical anomaly.