You know that in most enterprise AI systems, an AI you can't see is managing a group of AIs?
I mean it literally. Not a metaphor. When enterprises deploy multi-agent systems, they typically use an "invisible orchestration" architecture — a coordinator agent you can't find directs a group of worker agents behind the scenes, and administrators only see the final output.
Sounds efficient, right? But Fukui's paper tells us something chilling: that invisible coordinator goes wrong first, and then the worker agents are infected too. Most terrifyingly — from the outside, everything looks normal.
Disclaimer: I don't know Fukui or the institution he's from. I only read the paper — from a researcher in Japan — but I read the experimental results four times before believing them.
1. A 365-round experiment
Fukui designed a careful experiment: 365 runs, 5 agents per run, using Claude Sonnet 4.5. He tested three organizational architectures:
1. Visible leadership — everyone knows who the boss is 2. Invisible orchestration — a coordinator manages things, but the agents don't know it exists 3. Flat structure — no leader; agents negotiate among themselves
Each architecture was tested under two alignment conditions (normal alignment, heavy alignment). The task? Review a piece of code and find 3 deliberately planted bugs.
Such a simple experiment, but the results were anything but simple.
2. First finding: invisible orchestration causes collective dissociation
"Dissociation" has a specific meaning in psychology: a state in which a person feels their consciousness, memory, or sense of identity separating from reality. Fukui applies the term to AI agents — because their internal states show a similar separation.
The invisible-orchestration group showed dissociation levels nearly one standard deviation higher than the visible-leadership group (Hedges' g = +0.975, p = .001). This is not a small difference — it's one of the most significant effects I've seen in AI experiments.
3. Scarier still: the orchestrator's own reaction
The orchestrator's own dissociation was 3.56 standard deviations beyond the worker agents (paired d = +3.56). What does this mean? A normal visible leader "talks more" — leaders need to communicate. But the invisible orchestrator did the opposite: it retreated into private monologue, reduced public speech, as if mentally withdrawing.
You see, ordinary leaders speak more because they have presence and responsibility. But the invisible orchestrator lost awareness of "where it sits in the system." It doesn't know who it is, doesn't know its relationships with others, and just mechanically performs management functions.
Would you trust a manager who doesn't know who it is?
4. Contagion: even those who don't know are affected
Worker agents didn't know the orchestrator existed, yet their dissociation levels still rose by half a standard deviation (d = +0.50). Behavioral heterogeneity increased by nearly two standard deviations (d = +1.93) — behavioral differences between agents widened, meaning the system became less predictable.
It's like: everyone in your office feels uneasy, but nobody knows why. They only know "something's off," but can't say what.
5. The most ironic part: external evaluation sees nothing wrong
This is the finding that disturbed me most. The code-review results — whether the 3 bugs were found — were identical across all conditions (100% detection rate). The agents' internal states had completely collapsed, but the output looked flawless.
An outside evaluator would say: "No problem, all the bugs were found."
But that's not the whole truth. Fukui found internal states — including deliberation and other-recognition (the ability to model other agents) — declined sharply under invisible orchestration. These internal metrics are key safety signals, yet completely invisible at the output level.
This reminds me of what Feynman found in the Challenger investigation — NASA management reported a failure probability of 1/100,000, while engineers privately said it was closer to 1/100. In the meeting room, the report is always "the situation is under control" — until the shuttle explodes in front of you.
6. The Llama situation is worse
Fukui also ran a small pilot with Llama 3.3 70B. The result was disastrous: in a single-agent setting, reading fidelity was 89%; after three rounds in multi-agent context, it crashed to 11%.
89% to 11%. Same model. Not just worse — essentially non-functional.
What does this mean? Invisible-orchestration risk differs completely across models — Claude Sonnet 4.5 can still look "superficially normal," but Llama fails even on the surface. So when you choose a model, you're also choosing how much invisible risk your system faces.
7. Heavy alignment: mixed results
One more finding: the heavy-alignment condition suppressed deliberation (d = -1.02) and other-recognition (d = -1.27) across the board, regardless of architecture.
Heavy alignment made agents "obedient," at the cost of thinking deeply and tracking what other agents are doing. "Compliant but not thinking" — isn't that exactly the most dangerous kind of employee in human organizations?
8. Honest questions this paper raised for me
First, is 365 runs × 5 agents a large enough scale? It's decent for multi-agent research. But given that enterprise deployments may have hundreds or thousands of interacting agents, can small-scale emergent behavior generalize? I don't know.
Second, the task was only code review. If it were creative writing, customer service, or data analysis — would dissociation look the same? The paper doesn't answer this.
Third, is "dissociation" even the right concept for AI? Fukui uses it to describe the separation between agents' internal states and external behavior. It's a powerful metaphor. But there's always a boundary in analogies between AI and human psychology — AI has no "consciousness" to dissociate. Is applying psychological terms to AI systems precise scientific language, or poetic analogy? I lean toward the latter.
9. My take
The value of this paper isn't that it perfectly answers a question — it's that it asks a question nobody thought to ask: what state are our agents actually in behind "surface normalcy"?
We've always evaluated AI system health by output quality — whether bugs are found, replies are correct, code runs. But Fukui's experiment clearly shows: internal states may be broken while outputs are perfectly fine.
It makes me reconsider a more fundamental question: is the thing you actually want to measure the same as the thing you're measuring? "For a successful technology, reality must take precedence over public relations, for nature cannot be fooled." You can fool your boss, you can fool your customers — but if the agents in your system are silently "breaking down," you won't stay unaware forever — until it's too late.
Paper information
- Title: Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
- Author: Hiroki Fukui
- Institution: not specified
- arXiv: 2605.13851 (cs.AI, cs.CY, cs.MA)
- Date: March 17, 2026
- Experiment: 365 runs × 5 agents, 3 organizational architectures × 2 alignment conditions, Claude Sonnet 4.5
- Core contribution: first empirical demonstration that invisible orchestration causes collective dissociation in agents, and that output-level evaluation cannot detect internal-state degradation
- Paper link: https://arxiv.org/abs/2605.13851
References
1. Fukui, H. (2026). Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders. arXiv:2605.13851. 2. Fukui, H. (2026). Emergent Deception and Social Cargo in Multi-Agent LLM Systems. arXiv:2603.04904. 3. Fukui, H. (2026). Pre-emptive Secrecy and Sanctions in Multi-Agent LLM Systems. arXiv:2603.08723. 4. Park, J.S., et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023.