Why Smart Minds Self-Deceive: AI Scientific Reasoning and Confirmation Bias
> Paper: *FALSIFYBENCH: Evaluating Hypothesis-Driven Reasoning in LLMs* > arXiv: 2606.04751 > Authors: Leonardo Bertolazzi, Massimo W. Barberi, Maria Grazia > Posted: 2026-06-03
---
🕵️ Prelude: A Detective's Trap
Imagine a detective called to a murder scene. The victim clutches a button in his hand—apparently torn from the killer's coat. The detective immediately forms a hypothesis: *the killer is a person wearing this kind of coat*. He spends three days visiting every shop in town that sells such coats, investigating every buyer. On day five, he arrests a "suspect"—a young man missing a button.
But the real killer wore a completely different coat. The button? The victim had ripped it from his own jacket in his final moments, trying to leave a clue.
The detective is smart. His reasoning is sharp. His execution is relentless. Yet he made one fatal mistake: he spent the whole investigation looking for evidence that confirmed his hypothesis, instead of looking for evidence that could refute it.
This is *confirmation bias*—one of the most stubborn cognitive traps in human reasoning. On June 3, 2026, a team of researchers put this question to AI: when LLMs play the role of scientist, do they fall into the same trap?
---
🧪 1. The Wason 2-4-6 Task: A Classic Test of Scientific Thinking
1.1 A Simple but Treacherous Game
In 1960, British psychologist Peter Wason designed a deceptively simple experiment:
> The experimenter has a rule in mind—say, "three ascending numbers." He offers the triple *2-4-6*. The participant's job is to discover the rule by proposing triples; the experimenter replies "matches" or "doesn't match."
Most participants immediately guess: "The rule is *even numbers*!" They test 8-10-12 ("matches"), then 20-22-24 ("matches"), and feel confident. But the experimenter shakes his head. The real rule is the much simpler "three ascending numbers." 1-2-3 matches. 3-5-9 matches. 100-101-102 matches. Participants almost never test these, because they only test cases that would confirm their "even numbers" hypothesis.
They never ask: *what example, if it matched the rule, would prove my hypothesis wrong?* For instance, if they tested 1-2-3 and heard "matches," the "even numbers" hypothesis would collapse. But they almost never do.
1.2 Why This Task Cuts So Deep
The Wason 2-4-6 task became a psychology classic because it reveals a profound truth about human rationality:
We are biologically predisposed to seek confirming evidence, not falsifying evidence.
This is not stupidity. It may be an evolved cognitive shortcut: in daily survival, rapidly confirming a hypothesis ("Is there a lion in that grass?") is more useful than exhaustively testing it. But in scientific reasoning, this instinct becomes a fatal flaw. Karl Popper emphasized that the core of scientific method is falsification: a good theory must be *refutable*—there must exist possible observations that would prove it wrong.
---
🤖 2. FALSIFYBENCH: Putting AI in the Scientist's Seat
2.1 Porting the Classic Experiment to AI
The FALSIFYBENCH team brought Wason-style logic into the LLM era. They built a complete evaluation framework that probes multiple capabilities of scientific reasoning:
- Hypothesis generation: can the model propose reasonable hypotheses?
- Evidence gathering: can it design effective experiments to test them?
- Belief updating: does it revise its beliefs when evidence contradicts them?
- Falsification: does it actively seek disconfirming evidence?
- Bertolazzi, L., Barberi, M. W., & Grazia, M. (2026). *FALSIFYBENCH: Evaluating Hypothesis-Driven Reasoning in LLMs*. arXiv:2606.04751.
- Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. *Quarterly Journal of Experimental Psychology*, 12(3), 129-140.
- Popper, K. R. (1959). *The Logic of Scientific Discovery*. Hutchinson.
- Kuhn, T. S. (1962). *The Structure of Scientific Revolutions*. University of Chicago Press.
- Feynman, R. P. (1985). *Surely You're Joking, Mr. Feynman!*. W.W. Norton.
- Lakatos, I. (1978). *The Methodology of Scientific Research Programmes*. Cambridge University Press.
Twelve LLMs sat the exam, drawn from different families and scales, including reasoning models (such as the o1 series) and standard instruction-tuned models.
2.2 A Disquieting Result
The result is both predictable and troubling.
Reasoning-tuned models do outperform instruction-tuned ones at scientific reasoning—a relief, since it validates that the "reasoning" training direction is correct.
But then comes the harsh fact: no model approaches optimal performance. Even the best model falls far short of a trained human scientist. They can generate hypotheses and gather evidence to some degree, but at the most critical step—proactive falsification—almost every model fails.
2.3 Negative Testing: The Key Differentiator
The researchers ran a fine-grained *turn-level analysis*, examining not just outcomes but behavior on each interaction turn.
The finding is clear:
**The decisive difference between success and failure is whether the model performs *negative testing*.
Successful models proactively ask: "If my hypothesis is wrong, what evidence would I expect to see?"—and then search for that evidence. Failing models, like the humans in Wason's task, fall into the confirmation trap: they pile up evidence that supports the current hypothesis until they hit a wall and are forced to abandon it.
---
🧠 3. Deep Analysis: Why Is Falsification So Hard?
3.1 The Cognitive-Psychology View
From cognitive psychology, several deep mechanisms make falsification hard:
Sunk-cost effect Once cognitive resources have been invested in building a hypothesis, humans—and apparently LLMs—resist letting go. "I've been thinking about this so long, I should persist a bit more."
Cognitive dissonance When evidence contradicts a hypothesis, the result is psychological discomfort. The simplest relief? Ignore the contradiction, or reinterpret it.
Motivated reasoning** We don't merely want to *know the truth*; we want to *feel that we are right*. Falsification threatens that self-image.
These look like distinctly human mechanisms—yet LLMs exhibit similar behavior. The authors suggest this likely reflects *human bias in training data*: trained on vast corpora of human text, LLMs may have internalized human cognitive biases along with human knowledge.
3.2 The Logic Itself
Pure logic actually favors falsification. Popper's famous formulation: "A thousand white swans cannot prove that all swans are white, but a single black swan can falsify the proposition."
Logically, one counterexample outweighs a thousand positive cases. Psychologically, however, confirming a familiar pattern delivers a dopamine reward; discovering an error delivers cognitive pain. AI may face an analogous *reward-structure problem*: in training, generating "plausible" content is rewarded, while proactively questioning oneself may be penalized as "inconsistent" or "erroneous."
3.3 Navigating the Hypothesis Space
FALSIFYBENCH also reveals a problem of *hypothesis-space navigation*.
Picture scientific reasoning as a vast maze where each hypothesis is a path. Confirmation bias is: find a promising-looking path, then keep walking it until you hit a dead end. Falsification, by contrast, means: even on a promising path, actively search for evidence that the path is blocked. If it really is blocked, turn back early and try another.
The latter is plainly more efficient, but psychologically harder because it demands *giving up hope*. The researchers found that LLMs navigate hypothesis space in patterns strikingly similar to humans: they tend to *drill down* in local regions rather than explore globally.
---
🎭 4. Literary Reflections: Science, Detectives, and Self-Deception
4.1 Holmes's Method of Negative Reasoning
Sherlock Holmes famously said: "When you have eliminated the impossible, whatever remains, however improbable, must be the truth."
This is often misread as "gather more evidence." In fact, it is an art of *reverse thinking*: enumerate all possibilities, then systematically eliminate them. The method is, at its core, falsificationist. Holmes's brilliance lies not in superior deduction but in his refusal to commit prematurely to the first plausible hypothesis.
4.2 Feynan's First Principles
Richard Feynman may have been the 20th century's most gifted falsificationist. His habit: when someone proposed a theory, he would immediately think, "What experiment could prove this wrong?" Not out of contrarianism—he understood that a theory that can only be confirmed but never refuted is not science but faith.
In *Surely You're Joking, Mr. Feynman!*, he tells of attending a philosophy seminar where participants debated how to define science. Feynan's answer: "Science is the belief in the ignorance of experts." When you're unsure, you're unsure; when you're sure, it's because you have evidence.
This *epistemic humility*—perpetually holding the attitude "I might be wrong"—is the essence of the scientific spirit. FALSIFYBENCH shows that current AI is far from this ideal.
4.3 Orwell's Doublethink
In *1984*, George Orwell coined *Doublethink*: holding two contradictory beliefs simultaneously and accepting both. It is closely related to cognitive dissonance. Confirmation bias is a milder form of doublethink: we selectively attend to evidence that supports our beliefs and ignore contradictory evidence. Our belief system then *looks* consistent, while in reality it is full of contradictions.
If AI were to acquire this mild form of self-deception, the danger would be real: it would systematically err while wearing the appearance of rationality and consistency.
---
🔬 5. Echoes from Philosophy of Science: Popper, Kuhn, Lakatos
5.1 Popper's Falsificationism
In *The Logic of Scientific Discovery*, Karl Popper argued that the defining feature of a scientific theory is *refutability*. A good theory must risk being disproved. Science, on this view, is not the process of "confirming true theories" but of "eliminating false ones."
FALSIFYBENCH therefore tests not only reasoning ability but whether AI possesses the *scientific spirit* itself.
5.2 Kuhn's Paradigm Shifts
In *The Structure of Scientific Revolutions*, Thomas Kuhn refined Popper's picture. Most of the time, science is not "falsifying" but "puzzle-solving" within an existing paradigm. *Paradigm shifts*—abandoning the old framework for a new one—happen only when anomalies accumulate past a threshold.
FALSIFYBENCH's AIs mostly remain in the *puzzle-solving* stage. They can find confirming evidence within a given hypothesis, but struggle with true *paradigm shifts*: when the hypothesis itself is broken, abandoning it and seeking an entirely new framework.
5.3 Lakatos's Research Programmes
Imre Lakatos tried to reconcile Popper and Kuhn with the concept of a *research programme*: a "hard core" of inviolable basic assumptions, surrounded by a "protective belt" of adjustable auxiliary hypotheses. When faced with refutation, scientists adjust the belt, not the core.
FALSIFYBENCH suggests AI may need to draw a smarter line between belt and core. Some hypotheses should be flexibly adjusted; others should be defended to the hilt. The key is knowing which is which.
---
🌊 6. Real-World Echoes: From the Lab to Life
6.1 Confirmation Bias in Medical Diagnosis
Confirmation bias in medicine can be fatal. A physician who forms a diagnostic hypothesis too early ("this looks like pneumonia") and then attends only to supporting symptoms ("cough, fever") while ignoring contrary evidence ("atypical X-ray") risks serious misdiagnosis.
A good doctor asks: "If this isn't pneumonia, what evidence would prove me wrong?"—and actively seeks it.
If AI is to assist in diagnosis, it must acquire this habit of *proactive falsification*.
6.2 The Confirmation Trap in Investing
Investors often fall into confirmation bias: once they buy a stock, they see only good news and ignore warning signs.
Charlie Munger, Warren Buffett's partner, championed *inversion*: "All I want to know is where I'm going to die, so I'll never go there." Avoiding errors is, in essence, applying falsificationist thinking.
6.3 Self-Verification in Relationships
In social psychology, *self-verification theory* shows that people preferentially seek feedback that confirms their self-concept. A person with low self-esteem notices coldness in others but overlooks kindness; a narcissist notices praise but ignores criticism.
If an AI companion (such as Replika) is designed to *always affirm the user*, it may reinforce the user's cognitive biases rather than help them grow.
---
🔮 7. The Future: How to Teach AI Self-Doubt
7.1 A Revolution in Training Data
Teaching AI falsificationist thinking may require rethinking training corpora. Today's datasets are dominated by "correct content"—textbooks, encyclopedias, papers. Perhaps we need to add more historical records of *scientific error and correction*: how geocentrism was overturned by heliocentrism, how phlogiston theory was replaced by oxidation theory.
Showing AI *how great scientists erred and corrected themselves* may be more instructive than showing it only the correct answers.
7.2 Restructuring the Reward Function
In reinforcement learning, the reward function determines behavior. If the current reward signal only rewards "producing correct content," AI will drift toward confirmation bias. Add a reward for *proactively discovering one's own errors*, and AI may develop a healthier capacity for self-doubt.
7.3 Multi-Agent Debate
A promising approach: let multiple AIs debate. Each proposes a different hypothesis, then challenges the others' evidence. This kind of *adversarial verification* could simulate the functions of the scientific community—peer review, replication, open debate.
---
📚 References
*Auto-collected and interpreted on 2026-06-05*