Why Shared Whiteboards Make Small AI Agents Hallucinate More
This post presents a detailed walkthrough of the paper 'Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents' by Yunpeng Zhou (independent researcher), arXiv:2605.31354 (2026-05-29).
Paper at a glance
- Core finding: Shared workspaces in modular vision agents using low-capacity models (4B–8B) do not suppress hallucination—they *amplify* it.
- Two dominant failure modes: Noise Reinforcement and Policy Collapse.
- Counterintuitive result: More compute without explicit verification can correlate *negatively* with performance.
- Conclusion: The bottleneck is not reasoning depth but communication fidelity.
- Read — retrieve existing notes from the shared workspace
- Write — add new information based on observations and prior notes
- Verify — confirm reliability of written information
- Perceptual noise: visual misreads (small fonts, dense layouts, abbreviations in business documents).
- Reasoning noise: over-inference (e.g., 'sales up → profit up'), more dangerous because it wears a logical wrapper.
- Single-author paper; model coverage limited to 4B–8B; larger models untested.
- Experiments confined to DocVQA (deterministic answers); open-ended tasks unexplored.
- Verification itself costs ~25–30% extra compute.
- No comparison with human-team collaboration dynamics.
- Missing ablations quantifying each failure mode's contribution and their interactions.
- Mitigation strategies remain largely prescriptive rather than experimentally validated.
The shared-whiteboard metaphor
Imagine five interns with limited memory collaborating on a 50-page document via a shared whiteboard. Each writes unverified partial findings (with mistakes—misread decimal points, wrong tables); a synthesizer combines them into a self-consistent but false narrative. Errors from multiple agents stack, amplify, and mutually corroborate, producing a worse outcome than any individual working alone. This mirrors the paper's core phenomenon: hallucinations cross-infect in multi-agent collaboration.
CoSee: an audit framework for collaboration
CoSee decomposes modular agent collaboration into three primitive actions:
CoSee makes the previously black-box process explicit and traceable: which agent wrote what, who read it, whether it was ever verified. Experiments cover three DocVQA benchmarks (multi-page documents, charts, webpage screenshots), where no single agent can cover the full task, making collaboration necessary.
Failure mode 1: Noise Reinforcement
An agent writes an unverified claim (e.g., misreading '3.2%' as '32%'). Downstream agents treat it as fact, cite it, and even produce elaborate 'analyses' around it—granting the error false authority. Later agents trust the *prior agents' trust* rather than the original evidence. Error rates grow super-linearly as agent count rises from 2 to 5.
Two noise types are distinguished:
CoSee also documents a 'beautification effect': the more often an error is cited, the more credible it appears—a 'repetition-as-truth' bias.
Failure mode 2: Policy Collapse
As shared-workspace content accumulates, agents shift from careful analysis to perfunctory, under-specified answers—a cognitive shortcut under information overload. Quantitatively, a 4B model working alone averaged 47 tokens and 3.2 factual statements per answer; in collaboration, 19 tokens and 1.1. This resembles the human bystander effect: distributed responsibility.
A controlled experiment found that explicit role assignment performed worse than unassigned free-form collaboration. 'Verifier'-role agents over-relied on their role identity and produced many 'verified' labels with little genuine cross-checking—challenging role-play prompting as a sufficient verification mechanism.
The more-compute-is-worse paradox
Without explicit verification, raising collaboration rounds from 2 to 5 (~150% more FLOPs) *lowered* accuracy from 61% to 53%:
| Config | Rounds | Compute (GFLOPs) | Accuracy | |--------|--------|------------------|----------| | No verification | 2 | 12.4 | 61% | | No verification | 3 | 18.6 | 58% | | No verification | 5 | 31.0 | 53% | | Explicit verification | 2 | 15.8 | 64% | | Explicit verification | 3 | 23.7 | 71% | | Explicit verification | 5 | 39.5 | 78% |
Mechanisms behind the degradation:
1. Cumulative contamination: ~15–20% of new notes per round contain errors; by round 5, ~40% of notes hold detectable errors. 2. Over-complexified reasoning chains: low-capacity models lose track over long histories.
Only with explicit verification does the cost-accuracy Pareto frontier regain a normal upward trend.
Conclusion: the bottleneck is communication fidelity
For resource-constrained agents, single-agent intelligence cannot compensate for unreliable inter-agent channels. Communication fidelity spans information, structural, verification, and priority fidelity—all fail in 4B–8B models. Suggested (but not deeply tested) directions: structured shared spaces with provenance/confidence labels, redundant cross-verification, and information decay (half-life) for stale notes.
Limitations
References
1. Yao et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629 2. Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366 3. Li et al. (2023). Evaluating Object Hallucination in Large Vision-Language Models. EMNLP 2023. arXiv:2305.10355 4. Wei et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. arXiv:2201.11903 5. Kahneman, D. (2011). *Thinking, Fast and Slow*. Farrar, Straus and Giroux.