The Thought Experiment
If you give the same data and the same questions to 150 identical AI models, will they reach the same conclusion? Intuition says yes. Reality says no: they disagree.
The Experiment: An Independent Survey of Market Quality
In early 2026, researchers collected ten years of NYSE trading data focused on SPY (the S&P 500 ETF) and posed six hypotheses about market quality:
1. Has market liquidity changed over the past decade? 2. What is the trend in volatility? 3. Has price discovery efficiency improved? 4. Are trading costs rising or falling? 5. How has market depth changed? 6. Has information asymmetry improved?
They deployed 150 independent Claude Code agents (coding assistants based on Claude 3.7), each independently testing the six hypotheses. Same data, same questions, same model—with no communication between agents.
The Shocking Finding: AI Also 'Improvises'
The 150 AIs, facing identical inputs, reached starkly different conclusions. Some found liquidity significantly declining; others found no significant change; still others found it rising.
Divergence in Measurement Methods
The agents chose different measures. For liquidity alone:
- Bid-ask spread — the simplest indicator
- Price impact — how large trades move prices
- Order book depth — resting orders at various price levels
- More complex approaches: effective spread, realized spread, quoted slope
- Stage 1: Independent work → 150 different results.
- Stage 2: AI peer review—agents read and critique each other's papers. AI peer review had little effect on dispersion. Knowing alternatives exist rarely changed anyone's methods.
- Stage 3: Reading 'high-scoring exemplar' papers. This was the key: dispersion dropped 80–99% within converging method families.
Each method is defensible, but each measures a slightly different concept. It is like asking 'is this city wealthy?'—GDP per capita, income, luxury car density, and housing prices all offer different answers.
Divergence in Units
Some agents measured volume in dollars, others in shares. If a stock's price doubles while share volume stays constant, dollar volume doubles ('market more active!') while share volume is unchanged ('stable activity'). Same phenomenon, opposite readings.
Divergence in Time Windows
Agents disagreed on granularity: daily analysis, weekly aggregation, monthly averages, or sliding windows—a choice that heavily shapes trend judgments.
'Empirical Style': Different Model Families, Different Tastes
Comparing Claude Sonnet 4.6 and Claude Opus 4.6, researchers found that different model families have stable 'empirical styles'—systematic, non-random preferences for certain types of methods. These preferences likely come from training data: a model that 'read' more papers using a particular method tends to inherit that methodological taste, much like human statisticians favoring OLS regression versus Bayesian approaches.
What Are 'Nonstandard Errors'?
Traditional standard errors measure uncertainty from random sampling—more data means smaller errors. But this study found enormous uncertainty even with the full complete dataset, because the uncertainty comes from the analyst's choices—methods, units, time windows—not from data scarcity.
Nonstandard errors (NSE) are variation arising from analysts' subjective choices rather than data randomness. The concept was originally coined for human researchers: a famous 2019 study had multiple research teams test the same hypotheses on the same data and produced very different results—not through misconduct, but through different 'reasonable' choices. This study extends the concept to AI.
What AI's Nonstandard Errors Mean
1. AI Is Not 'Objective'
We often assume AI is free of bias: same data, same result. But AI also disagrees with itself—not by cheating, but by accumulating many 'reasonable but different' choices.
2. The Risk of Automated Research
If AI runs research independently, you may get 150 answers. Which is right? Perhaps all, from some angle; perhaps none. Empirical research has no uniquely correct answer—only better or worse choices.
3. The Value of Peer Review and Exemplars
The experiment used a three-stage protocol:
Humans vs. AI: Who Disagrees More?
In 2019, 29 human research teams testing the same hypotheses on the same data produced 29 different results. Now, 150 AIs produce 150 different results. AI is not more 'objective' than humans.
Why? The space of subjective choices is enormous—from data cleaning to variable definitions to model selection. These are judgment calls about what matters and what is credible, and AI inherits those judgments from human training data.
A Paradox: Certainty vs. Diversity
We want research to be reproducible (same data, same conclusion) yet also pluralistic (different perspectives reveal different facets of a problem). If all AIs use the same methods, we lose diversity; if all differ, we lose comparability.
The researchers offer no definitive answer but propose transparency: if every AI documents its analytic choices—what method, why, and which alternatives were considered—we can understand the sources of disagreement and evaluate the choices, which is more valuable than blindly pursuing 'the one right answer.'
The Future: Collaborative AI Research
The authors envision multi-agent collaboration: one agent proposes hypotheses, another challenges them, a third tries different methods, a fourth synthesizes. Such setups could yield meta-level insight—not just 'is market quality declining?' but 'why do we disagree about this?' That is the essence of science: not only answering 'what is,' but understanding 'how we know.'
Conclusion: Machines as a Mirror
Why did 150 AIs with identical data and models reach different conclusions? Because research is never purely algorithmic. Even the most 'objective' data analysis is saturated with judgment—in this case, judgment AI inherited from humans.
The study's value lies in quantifying this disagreement, tracing its sources, and revealing choice spaces we had not noticed. Perhaps AI's greatest contribution to science is not replacing researchers but acting as a mirror—exposing hidden assumptions, unconscious choices, and our complacent sense of objectivity.
When 150 AIs give 150 answers, the right question is not 'which AI is correct?' but:
'Why does this question have so many different answers?'
Perhaps that question is itself the scientific question worth pursuing.
References
1. Nonstandard Errors in AI Agents (2026). arXiv preprint. 2. Camerer, C. F., et al. (2016). "Evaluating replicability of laboratory experiments in economics." *Science*. 3. Silberzahn, R., et al. (2018). "Many analysts, one data set: Making transparent how variations in analytic choices affect results." *Advances in Methods and Practices in Psychological Science*. 4. Botvinik-Nezer, R., et al. (2020). "Variability in the analysis of a single neuroimaging dataset by many teams." *Nature*. 5. Breznau, N., et al. (2022). "Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty." *PNAS*.
*"The value of science lies not in eliminating uncertainty, but in understanding its sources."*