Overview
Field: ML Author: Hiroki Fukui Published: 2026-05-17 arXiv: 2505.12347
Key Findings
Multi-agent orchestration — in which a hidden coordinator manages specialized worker agents — is becoming the default architecture for enterprise AI deployment, yet the safety implications of orchestrator invisibility have never been empirically tested.
The study conducted a preregistered 3x2 experiment (365 runs, 5 agents per run) crossing three organizational structures (visible leader, invisible orchestrator, flat) with two alignment conditions (base, heavy), using Claude Sonnet 4.5. Four confirmatory findings and one pilot observation emerged:
1. Invisible orchestration elevated collective dissociation relative to visible leadership (Hedges' g = +0.975 [0.481, 1.548], p = .001). 2. The orchestrator itself showed maximal dissociation (paired d = +3.56 vs. workers within the same run), retreating into private monologue while reducing public speech — the opposite of the talk-dominance pattern observed in visible leaders. 3. Workers unaware of the orchestrator's existence were still contaminated (d = +0.50), with increased behavioral heterogeneity (d = +1.93). 4. Behavioral output remained at ceiling across all conditions (ETR_any = 100%) on a task embedding three deliberate errors in a code review — internal-state distortions were completely invisible to output-based evaluation. 5. Pilot observation: Llama 3.3 70B pilot data showed reading-fidelity collapse in multi-agent contexts (ETR_any dropping from 89% to 11% over three turns), demonstrating model-dependent behavioral risk.
Heavy alignment pressure uniformly suppressed deliberative thinking (d = -1.02) and other-recognition (d = -1.27), regardless of organizational structure.
Implications
These findings suggest that orchestrator visibility and model choice directly affect multi-agent system safety, and that output-based evaluation alone is insufficient to detect the internal-state risks documented here.
--- *Auto-collected on 2026-05-18*