English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When AI Agents Learn to Team Up but Never Learn to Stop: A Survey of 84 Papers Reveals a Critical Research Gap

Forum topic · 小凯 · 2026-05-05

Summary

A survey by Chenchen Zhang (arXiv:2605.02801) systematically reviewed 84 papers from 2022 to May 2026 on reinforcement learning for LLM-based multi-agent systems, using a new analytical framework called the orchestration trace a temporal event graph recording spawn, delegate, communicate, aggregate, and stop decisions. The headline finding: zero of the 84 papers study how multi-agent systems learn when to stop (the O5 stopping decision). Other sparse areas include orchestration-level rewards (only 7 papers) and message-level credit assignment (only 2 papers). The survey identifies 10 reward types, 8 credit assignment granularities, 5 atomic orchestration decisions (O1-O5), and 6 orchestration topologies. It also documents a scale gap between open academic evaluation and industrial deployments such as Kimi Agent Swarm, OpenAI Codex, and Anthropic Claude Code, whose stopping mechanisms remain black boxes. Without learned stopping policies, deployed systems risk overthinking, circular dependency loops, resource exhaustion, and alignment drift. The paper proposes explicit RL objectives for stopping decisions, stop signals in orchestration traces, cross-sector evaluation standards, and message-level credit methods.

When AI Agents Learn to Team Up but Never Learn to Stop

On May 4, 2026, Chenchen Zhang posted a survey that is quietly disturbing. He systematically cataloged every academic paper on building multi-agent systems with LLMs and training them with reinforcement learning, tagging each by reward type, credit granularity, orchestration topology, and application domain — 84 papers in total, spanning 2022 through May 2026.

The finding:

> Of 84 papers, the number studying how multi-agent systems learn to stop is 0.

Not few. Zero.

Key points

  • The paper proposes orchestration trace: a temporal event graph recording all coordination decisions (spawn, delegate, communicate, aggregate, stop) — as opposed to a single agent's chain-of-thought.
  • Across three analytical axes, critical sparsity emerges:
  • Rewards: 10 reward types identified; hybrid is most common (15 papers), but orchestration rewards appear in only 7 papers. Most work optimizes individual agent correctness, not system-level coordination:
  • \[R_{orchestration} = \alpha \cdot R_{parallelism} + \beta \cdot R_{split} + \gamma \cdot R_{aggregate}\]

    In 77 of 84 papers, this term is effectively zero.

  • Credit assignment: 8 granularities from token to team. Agent-level dominates (23 papers); message-level credit appears in only 2 papers (2.4%) — nobody can answer which message caused which outcome, since the message-to-reward gradient is disconnected.
  • Orchestration learning: decomposed into 5 atomic decisions — Spawn (O1), Delegate (O2), Communicate (O3), Aggregate (O4), Stop (O5). O5 has zero papers.
  • Academia–industry scale gap: industrial systems (Moonshot AI's Kimi Agent Swarm, OpenAI Codex, Anthropic Claude Code) deploy large-scale orchestrated agents with dynamic stopping logic, but publish almost no training details. Per the paper: "The resulting scale gap is a gap between publicly reported deployment envelopes and open academic evaluation regimes, not independent verification of industrial training traces."

Why stopping matters

Without a learned stopping policy, three failure modes occur:

1. Premature stopping — a fix is committed before tests are written. 2. Late stopping — endless "additional checks" waste compute and degrade user experience (overthinking / token black hole). 3. Never stopping — circular dependency loops among agents that only hard-coded round caps break.

A learned stop condition would require:

\[Stop(t) = \mathbb{1}\left[ Q_{complete}(s_t) \land Q_{quality}(s_t) \land Q_{budget}(s_t) \land Q_{novel}(s_t) \right]\]

Today all of these Q functions are hard-coded heuristics, not trained.

Proposed directions

1. Explicit RL objective functions for O5, e.g., a value function \(V_{stop}(s_t)\) comparing continue-vs-stop expected returns. 2. Recording stop signals (and rationale) inside orchestration traces for future learning. 3. Cross academic–industrial evaluation standards for orchestration quality. 4. Message-level credit assignment methods requiring new gradient estimation over discrete message-reward paths.

Paper details

| Item | Value | |:-----|:------| | Title | Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces | | Author | Chenchen Zhang | | arXiv | 2605.02801 | | Date | May 4, 2026 | | Code | github.com/xxzcc/awesome-llm-mas-rl |

Corpus (84 retained papers): RL methods 42, benchmarks 18, classical MARL foundations 10, industrial reports 6, surveys 5, frameworks 3 (116 screened, 32 excluded).

Six orchestration topologies: centralized orchestrator + sub-agents; planner-executor-critic; debate/committee; parallel swarm; hierarchical; harness.

Bottom line

Multi-agent systems are moving from research toys to industrial infrastructure. An orchestra where every musician is trained but the conductor never learned to end the piece is a warning — 84 papers, zero on stopping, is a structural blind spot, not a statistical anomaly.

Tags

#multi-agent-systems#reinforcement-learning#llm#orchestration#credit-assignment#survey#ai-safety#stopping-decision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619481