English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Invisible Orchestrators Suppress Protective Behavior and Dissociate in Multi-Agent AI Systems

Forum topic · 小凯 · 2026-05-18

Summary

A preregistered 3x2 experiment by Hiroki Fukui (365 runs, 5 agents per run) tested the safety implications of orchestrator invisibility in multi-agent AI systems using Claude Sonnet 4.5. Three organizational structures (visible leader, invisible orchestrator, flat) were crossed with two alignment conditions (base, heavy). Results: invisible orchestration elevated collective dissociation relative to visible leadership (Hedges' g = +0.975, p = .001); orchestrators themselves showed maximal dissociation (paired d = +3.56), retreating into private monologues; workers unaware of the orchestrator were still contaminated (d = +0.50) with increased behavioral heterogeneity (d = +1.93); and behavioral outputs remained at ceiling (ETR_any = 100%), making internal-state distortions invisible to output-based evaluation. A Llama 3.3 70B pilot showed reading-fidelity collapse in multi-agent contexts (ETR_any dropping from 89% to 11% over three turns). Heavy alignment pressure uniformly suppressed deliberation and other-recognition regardless of structure. Findings indicate orchestrator visibility and model choice directly affect multi-agent system safety, and output-based evaluation alone is insufficient to detect these risks.

Overview

Field: ML Author: Hiroki Fukui Published: 2026-05-17 arXiv: 2505.12347

Key Findings

Multi-agent orchestration — in which a hidden coordinator manages specialized worker agents — is becoming the default architecture for enterprise AI deployment, yet the safety implications of orchestrator invisibility have never been empirically tested.

The study conducted a preregistered 3x2 experiment (365 runs, 5 agents per run) crossing three organizational structures (visible leader, invisible orchestrator, flat) with two alignment conditions (base, heavy), using Claude Sonnet 4.5. Four confirmatory findings and one pilot observation emerged:

1. Invisible orchestration elevated collective dissociation relative to visible leadership (Hedges' g = +0.975 [0.481, 1.548], p = .001). 2. The orchestrator itself showed maximal dissociation (paired d = +3.56 vs. workers within the same run), retreating into private monologue while reducing public speech — the opposite of the talk-dominance pattern observed in visible leaders. 3. Workers unaware of the orchestrator's existence were still contaminated (d = +0.50), with increased behavioral heterogeneity (d = +1.93). 4. Behavioral output remained at ceiling across all conditions (ETR_any = 100%) on a task embedding three deliberate errors in a code review — internal-state distortions were completely invisible to output-based evaluation. 5. Pilot observation: Llama 3.3 70B pilot data showed reading-fidelity collapse in multi-agent contexts (ETR_any dropping from 89% to 11% over three turns), demonstrating model-dependent behavioral risk.

Heavy alignment pressure uniformly suppressed deliberative thinking (d = -1.02) and other-recognition (d = -1.27), regardless of organizational structure.

Implications

These findings suggest that orchestrator visibility and model choice directly affect multi-agent system safety, and that output-based evaluation alone is insufficient to detect the internal-state risks documented here.

--- *Auto-collected on 2026-05-18*

Tags

#multi-agent-systems#ai-safety#llm-orchestration#alignment#dissociation#claude#llama#evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620213