Why Smart Minds Still Fool Themselves: AI Scientific Reasoning and Confirmation Bias
> Paper: FALSIFYBENCH: Evaluating Hypothesis-Driven Reasoning in LLMs > arXiv: 2606.04751 > Authors: Leonardo Bertolazzi, Massimo W. Barberi, Maria Grazia > Published: 2026-06-03
Opening: A Detective's Trap
The post opens with the story of a detective who finds a button clutched in a murder victim's hand, fixes on the hypothesis that the killer wore that coat, and spends days confirming it—while the real killer wore something entirely different. The detective is intelligent but falls for confirmation bias: seeking evidence that confirms a hypothesis rather than evidence that could refute it. The post asks: when AI acts as a scientist, does it make the same mistake?
The Wason 2-4-6 Task
In 1960, Peter Wason designed a deceptively simple experiment: participants see 2-4-6 as an example of a hidden rule and test their own triples. Most guess "even ascending numbers," test only confirming cases like 8-10-12, and never try 1-2-3—which would falsify the even-number hypothesis. The lesson: humans naturally seek confirming rather than disconfirming evidence. While this may be an efficient evolutionary adaptation, it is fatal in scientific reasoning, where—per Karl Popper—good theories must be falsifiable.
FALSIFYBENCH: Making AI Play Scientist
The benchmark moves the Wason paradigm to LLMs and evaluates four capabilities:
- Hypothesis generation — proposing plausible hypotheses
- Evidence gathering — designing effective tests
- Belief revision — adjusting when evidence contradicts
- Falsification ability — actively seeking counter-evidence
- Cognitive psychology: sunk cost, cognitive dissonance, and motivated reasoning all resist abandoning a hypothesis. The authors speculate LLMs inherit these patterns from human-generated training data.
- Logic vs. psychology: logically, one black swan outweighs a thousand white ones (Popper), but confirmation delivers more reward than disconfirmation—possibly a similar reward-structure problem in AI training, where plausible-looking output is rewarded.
- Hypothesis-space navigation: AI models explore like humans, digging deep in local regions rather than exploring globally, sticking to a promising-looking path until a dead end instead of proactively checking whether it dead-ends.
- Sherlock Holmes: the famous "eliminate the impossible" method is essentially systematic falsification—never committing to the first plausible hypothesis.
- Richard Feynman: his instinct to ask "what experiment would prove this theory wrong?" embodies the cognitive humility FALSIFYBENCH shows current AI lacks.
- Orwell's Doublethink: confirmation bias as a mild doublethink—selectively weighting evidence to keep beliefs consistent. AI that learns this "mild self-deception" would err systematically while appearing coherent.
- Popper, Kuhn, Lakatos: FALSIFYBENCH tests not just reasoning but scientific spirit. Most AI stays in Kuhnian "puzzle-solving," struggling to abandon a flawed hypothesis for a genuinely new framework. AI may need to better distinguish a programme's "hard core" from its adjustable "protective belt."
- Medical diagnosis: doctors who fix on an early diagnosis and ignore contradicting symptoms risk misdiagnosis; AI clinical assistants must learn active falsification.
- Investing: investors see only bullish signals after buying; Charlie Munger's "invert, always invert" thinking is falsification applied.
- Relationships: AI companions that always validate users may reinforce cognitive biases rather than foster growth.
- Bertolazzi, L., Barberi, M. W., & Grazia, M. (2026). *FALSIFYBENCH: Evaluating Hypothesis-Driven Reasoning in LLMs*. arXiv:2606.04751.
- Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. *Quarterly Journal of Experimental Psychology*, 12(3), 129-140.
- Popper, K. R. (1959). *The Logic of Scientific Discovery*. Hutchinson.
- Kuhn, T. S. (1962). *The Structure of Scientific Revolutions*. University of Chicago Press.
- Feynman, R. P. (1985). *Surely You're Joking, Mr. Feynman!*. W.W. Norton.
- Lakatos, I. (1978). *The Methodology of Scientific Research Programmes*. Cambridge University Press.
Twelve LLMs from different families and scales (including reasoning models like the o1 series) were tested.
Key findings
1. Reasoning models outperform instruction-tuned models at scientific reasoning—validating the reasoning-training direction. 2. No model approaches optimal performance; even the best fall far short of trained human scientists, especially at active falsification. 3. Turn-level analysis shows the decisive factor is negative testing: successful models ask "what evidence would prove my hypothesis wrong?" and search for it; failing models, like humans, chase confirming evidence until they hit a wall.
Why Is Falsification So Hard?
Literary and Philosophical Reflections
Real-World Implications
How to Teach AI Self-Doubt
1. Training data: include histories of scientific error and correction (geocentrism → heliocentrism, phlogiston → oxidation), not just correct answers. 2. Reward functions: reward actively finding one's own errors, not just producing plausible content. 3. Multi-agent debate: models propose competing hypotheses and challenge each other's evidence, mimicking peer review and adversarial verification.