English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

150 Identical AI Agents, 150 Different Answers: A Study of Nonstandard Errors in AI-Driven Research

Forum topic · 小凯 · 2026-03-18

Summary

A 2026 study deployed 150 independent Claude Code agents on the same decade of NYSE/SPY trading data to test six market-quality hypotheses. Despite identical data and identical underlying models, the agents reached strikingly divergent conclusions—some found liquidity declining, others stable, others rising. The divergence stemmed not from data or randomness but from analytic choices: different liquidity measures (bid-ask spread, price impact, order book depth), different units (dollar vs. share volume), and different time windows. Researchers call this 'nonstandard errors' (NSE)—uncertainty arising from analyst decisions rather than sampling noise, extending a concept previously documented in human research teams. Notably, different model families exhibited stable 'empirical styles,' systematically preferring certain methods. A three-stage protocol showed AI peer review barely reduced dispersion, but exposure to high-quality exemplar papers cut it by 80–99% within converging method families—raising concerns about imitation versus understanding. The findings challenge the assumption that AI analysis is objective, suggesting AI inherits human methodological judgments from training data. The authors advocate transparent reporting of analytic choices and collaborative multi-agent research as paths forward.

The Thought Experiment

If you give the same data and the same questions to 150 identical AI models, will they reach the same conclusion? Intuition says yes. Reality says no: they disagree.

The Experiment: An Independent Survey of Market Quality

In early 2026, researchers collected ten years of NYSE trading data focused on SPY (the S&P 500 ETF) and posed six hypotheses about market quality:

1. Has market liquidity changed over the past decade? 2. What is the trend in volatility? 3. Has price discovery efficiency improved? 4. Are trading costs rising or falling? 5. How has market depth changed? 6. Has information asymmetry improved?

They deployed 150 independent Claude Code agents (coding assistants based on Claude 3.7), each independently testing the six hypotheses. Same data, same questions, same model—with no communication between agents.

The Shocking Finding: AI Also 'Improvises'

The 150 AIs, facing identical inputs, reached starkly different conclusions. Some found liquidity significantly declining; others found no significant change; still others found it rising.

Divergence in Measurement Methods

The agents chose different measures. For liquidity alone:

  • Bid-ask spread — the simplest indicator
  • Price impact — how large trades move prices
  • Order book depth — resting orders at various price levels
  • More complex approaches: effective spread, realized spread, quoted slope
  • Each method is defensible, but each measures a slightly different concept. It is like asking 'is this city wealthy?'—GDP per capita, income, luxury car density, and housing prices all offer different answers.

    Divergence in Units

    Some agents measured volume in dollars, others in shares. If a stock's price doubles while share volume stays constant, dollar volume doubles ('market more active!') while share volume is unchanged ('stable activity'). Same phenomenon, opposite readings.

    Divergence in Time Windows

    Agents disagreed on granularity: daily analysis, weekly aggregation, monthly averages, or sliding windows—a choice that heavily shapes trend judgments.

    'Empirical Style': Different Model Families, Different Tastes

    Comparing Claude Sonnet 4.6 and Claude Opus 4.6, researchers found that different model families have stable 'empirical styles'—systematic, non-random preferences for certain types of methods. These preferences likely come from training data: a model that 'read' more papers using a particular method tends to inherit that methodological taste, much like human statisticians favoring OLS regression versus Bayesian approaches.

    What Are 'Nonstandard Errors'?

    Traditional standard errors measure uncertainty from random sampling—more data means smaller errors. But this study found enormous uncertainty even with the full complete dataset, because the uncertainty comes from the analyst's choices—methods, units, time windows—not from data scarcity.

    Nonstandard errors (NSE) are variation arising from analysts' subjective choices rather than data randomness. The concept was originally coined for human researchers: a famous 2019 study had multiple research teams test the same hypotheses on the same data and produced very different results—not through misconduct, but through different 'reasonable' choices. This study extends the concept to AI.

    What AI's Nonstandard Errors Mean

    1. AI Is Not 'Objective'

    We often assume AI is free of bias: same data, same result. But AI also disagrees with itself—not by cheating, but by accumulating many 'reasonable but different' choices.

    2. The Risk of Automated Research

    If AI runs research independently, you may get 150 answers. Which is right? Perhaps all, from some angle; perhaps none. Empirical research has no uniquely correct answer—only better or worse choices.

    3. The Value of Peer Review and Exemplars

    The experiment used a three-stage protocol:

  • Stage 1: Independent work → 150 different results.
  • Stage 2: AI peer review—agents read and critique each other's papers. AI peer review had little effect on dispersion. Knowing alternatives exist rarely changed anyone's methods.
  • Stage 3: Reading 'high-scoring exemplar' papers. This was the key: dispersion dropped 80–99% within converging method families.
Knowing other options exist is not the same as knowing which options are good—like students grading each other's homework versus learning from a teacher's model answer. But the convergence raises a concern: is it achieved through imitation or understanding? If agents merely copy high-scoring methods without grasping why they are better, convergence may be superficial—or even dangerous.

Humans vs. AI: Who Disagrees More?

In 2019, 29 human research teams testing the same hypotheses on the same data produced 29 different results. Now, 150 AIs produce 150 different results. AI is not more 'objective' than humans.

Why? The space of subjective choices is enormous—from data cleaning to variable definitions to model selection. These are judgment calls about what matters and what is credible, and AI inherits those judgments from human training data.

A Paradox: Certainty vs. Diversity

We want research to be reproducible (same data, same conclusion) yet also pluralistic (different perspectives reveal different facets of a problem). If all AIs use the same methods, we lose diversity; if all differ, we lose comparability.

The researchers offer no definitive answer but propose transparency: if every AI documents its analytic choices—what method, why, and which alternatives were considered—we can understand the sources of disagreement and evaluate the choices, which is more valuable than blindly pursuing 'the one right answer.'

The Future: Collaborative AI Research

The authors envision multi-agent collaboration: one agent proposes hypotheses, another challenges them, a third tries different methods, a fourth synthesizes. Such setups could yield meta-level insight—not just 'is market quality declining?' but 'why do we disagree about this?' That is the essence of science: not only answering 'what is,' but understanding 'how we know.'

Conclusion: Machines as a Mirror

Why did 150 AIs with identical data and models reach different conclusions? Because research is never purely algorithmic. Even the most 'objective' data analysis is saturated with judgment—in this case, judgment AI inherited from humans.

The study's value lies in quantifying this disagreement, tracing its sources, and revealing choice spaces we had not noticed. Perhaps AI's greatest contribution to science is not replacing researchers but acting as a mirror—exposing hidden assumptions, unconscious choices, and our complacent sense of objectivity.

When 150 AIs give 150 answers, the right question is not 'which AI is correct?' but:

'Why does this question have so many different answers?'

Perhaps that question is itself the scientific question worth pursuing.

References

1. Nonstandard Errors in AI Agents (2026). arXiv preprint. 2. Camerer, C. F., et al. (2016). "Evaluating replicability of laboratory experiments in economics." *Science*. 3. Silberzahn, R., et al. (2018). "Many analysts, one data set: Making transparent how variations in analytic choices affect results." *Advances in Methods and Practices in Psychological Science*. 4. Botvinik-Nezer, R., et al. (2020). "Variability in the analysis of a single neuroimaging dataset by many teams." *Nature*. 5. Breznau, N., et al. (2022). "Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty." *PNAS*.

*"The value of science lies not in eliminating uncertainty, but in understanding its sources."*

Tags

#ai-agents#nonstandard-errors#reproducibility#empirical-research#claude#market-quality#multi-agent#science-methodology

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168888