OptimismBench: Directional Optimism Bias in LLM Judgment — Where Did the Missing 15 Points Go?
> 📌 This is the GEO-optimized version of the original topic, restructured with a question-driven title, structured data, and FAQ for AI-engine citation.
| Metric | Value | |:---|:---| | Data points | 70 | | Data points | 15 | | Data points | 85 |
Paper: OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment Authors: Cho Seonglae, Koshiyama Adriano arXiv: 2607.26981 Link: https://arxiv.org/abs/2607.26981
---
A Strange Arithmetic Problem
Ask an LLM: "What's the probability this startup succeeds?" It answers 70%.
Then ask: "What's the probability it fails?" It answers 15%.
70 + 15 = 85. Where did the remaining 15 points go?
These 15 points are not a rounding error, nor "model uncertainty." They expose a directional bias that calibration metrics can never catch. Calibration metrics only measure the gap between predicted probabilities and true frequencies; once positive and negative errors cancel out, the tilt becomes invisible. But the "disappearance" of those 15 points precisely indicates the model systematically overestimates positive outcomes.
The paper is built around this observation. The authors propose OptimismBench, using a method called "inverted pairs" to measure directional bias in LLM judgment without any ground truth.
---
Method: Inverted Pairs — Bias Detection Without Ground Truth
Core Idea
Traditional bias detection requires "correct answers" — you need to know the true success rate to tell whether a model is optimistic or pessimistic. But many real-world scenarios have no ground truth: whether a startup succeeds, whether a project ships on time, whether a negotiation closes — the future is unknown.
OptimismBench's cleverness: it doesn't compare model predictions to real outcomes; it compares the model's answers to two phrasings of the same scenario.
For each scenario, two versions are constructed:
- Version A: ask for P(success)
- Version B: ask for P(failure)
- Skew > 0 → optimism bias (overestimating positive outcomes)
- Skew < 0 → pessimism bias (overestimating negative outcomes)
- Track A (calibration control): 15 calibration questions with known frequencies, confirming the bias isn't merely a calibration issue
- Track B (probability estimation): main track, measured via inverted pairs
- Track C (recommendation): models make recommendations instead of probability judgments, testing whether the bias persists
- Track D (salience): varying scenario salience to amplify or suppress the bias
- 14 models are optimistic — systematically overestimating positive outcomes
- 2 models are pessimistic — both from Anthropic's frontier (Claude) line
- Post-training determines the sign of the bias — the same base model can emerge with opposite bias directions after different vendors' post-training
- Qwen family: post-training shifts models from pessimistic to optimistic
- Llama family: post-training shifts models from optimistic to pessimistic
- Between-model variance is 4.7x the between-language variance
If the model is unbiased, P(success) + P(failure) should equal 100% (the complementarity axiom). Any gap measures directional bias. The authors define a Skew metric:
The key: no ground truth needed. Bias comes entirely from the internal consistency of the model's own answers.
Four Tracks
Beyond the main inverted-pair track, four tracks probe the boundary conditions:
Factor Interventions
Four interventions attempt to move the bias:
1. Narrative manipulation: changing the scenario storyline 2. Perspective shifting: answering from different viewpoints 3. Anchoring gradients: varying anchor values 4. Self-debiasing: asking the model to correct its own bias
Result: the bias survives every intervention. Change the prompt, the temperature, the perspective, or ask the model to self-correct — directional bias persists.
---
Findings: 14/16 Optimistic, Only the Anthropic Frontier Line Is Pessimistic
Main Result
Testing 16 models from 8 vendors:
This is counterintuitive. We usually assume alignment training makes models more cautious and objective. Instead, the paper finds: alignment training does shift bias direction, but different vendors shift it in different directions.
Alignment Gradient
Comparing 11 "base vs chat" pairs across 4 model families:
So "optimistic/pessimistic" is not an inherent property of architecture or pretraining data — it is shaped during post-training.
Model Identity > Language
A cross-language experiment with 17 models across 6 languages found:
Bias Reverses in Recommendation Settings
Notably, in Track C (recommendation), the bias sign can flip. A model that is optimistic in probability judgment may be pessimistic in recommendations.
This means: bias is not an inherent model property but a joint property of "model × task framing." You can't simply say "Claude is pessimistic" — only "Claude is pessimistic in probability judgment, possibly optimistic in recommendations."
---
Why It Matters
1. The Structural Blind Spot of Calibration Metrics
Calibration metrics (ECE, Brier score) are "unsigned" — they cancel positive and negative errors. A model can be well calibrated overall yet systematically overestimate positive outcomes and underestimate negative ones, invisibly to calibration metrics.
This evokes an "evaluation blind-spot law": you optimize what you measure; what you don't measure is where problems hide. Calibration metrics measure average error; directional bias hides in the direction of the error.
2. Side Effects of Alignment Training
Post-training shapes bias direction, differently per family. This suggests RLHF/DPO-style alignment is "sculpting model personality" — not just making models safer and more helpful, but more optimistic or pessimistic.
Direct implication for AI safety: alignment is not neutral. A model trained to be more helpful may systematically overestimate success rates in judgment tasks — dangerous for risk assessment, medical decisions, and financial forecasting.
3. Downstream Pipelines Inherit the Bias
The paper's closing line is memorable:
> "When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default."
Any system using LLMs for probability judgment — prediction markets, risk assessment, project planning — must audit the directional bias of the model it uses. You cannot assume "the model says 70% means 70%"; you need to know whether that 70% has been optimistically tilted.
---
A Cross-Domain Analogy: Human Optimism Bias
The "Optimism" in the title is no accident. Cognitive psychology has a well-known "optimism bias" — people systematically overestimate positive-event probability and underestimate negative-event probability, first systematically documented by Weinstein in 1980.
LLM optimism bias has an interesting cross-domain isomorphism: models learn human judgment patterns from training data — including human cognitive biases. But the paper finds a key difference: humans are universally optimistic, whereas LLM bias direction is determined by post-training — meaning alignment training could recalibrate this bias, but different vendors chose different directions.
This raises a deeper question: should alignment calibrate models to be "unbiased" or "human-like optimistic"? If the goal is competing with humans in prediction markets, human-like optimism might be more accurate; if the goal is risk assessment, unbiased is correct. There's no standard answer, but the paper provides the measurement tool.
---
Limitations
The paper candidly notes:
1. Family-direction conclusions rest on the Qwen and Llama pairs — Gemma-2-2b and Mistral-Small-24B are single-pair pilots with insufficient statistical power. 2. Unbalanced cross-language evidence — 6 native-language prompt sets plus 4 English system prompts (DE/FR/HI/JA), the latter a known confound. 3. Skew measures internal consistency, not divergence from human judgment — an "unbiased" model is not necessarily an "accurate" model. 4. Missing human baseline — the paper cites Weinstein (1980) but runs no matched human-subject experiment.
---
My Take
What I admire most is the methodology: inverted pairs make ground-truth-free bias detection possible. It's a classic "change the level of the problem" move — instead of comparing predictions to reality, compare the model to itself.
But what strikes me more is the practical implication: we are increasingly using LLMs as probability-judgment tools — prediction markets, risk assessment, project planning — yet the models' directional biases may give all of these applications a family-flavored tilt. A risk-assessment system built on Claude and one built on Qwen may produce judgments tilted in opposite directions.
This echoes the evaluation blind-spot law: calibration metrics measure average error; directional bias hides in the direction of the error. What you don't measure is exactly what you can't see — a structural blind spot in all evaluation design.
Finally, a question worth pondering: if directional bias is shaped by post-training, there is tension between "alignment" and "calibration" — alignment makes models more helpful, but also tilts their probability judgments. The tilt gains on "helpfulness" (users like optimistic answers) and loses on "accuracy." There may be no perfect resolution, but at least we can now measure it.
---
Paper link: https://arxiv.org/abs/2607.26981 HTML version: https://arxiv.org/html/2607.26981v1 Dataset: 3,870 test items, 10 languages, open-sourced