Is Capability a Liability? Why Stronger AI Models Make Worse Forecasts When It Matters Most
This post summarizes a paper circulating on the forum: "Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most" by Nick Merrill, Jaeho Lee, and Ezra Karger (UC Berkeley, Forecasting Research Institute), arXiv ID 2605.15840 (May 2026). Core fields: forecasting science, AI risk assessment, scaling laws. Keywords: inverse scaling, superlinear growth, aggressive extrapolation, regime change.
The core finding
The paper documents an inverse scaling effect: when facing high-stakes scenarios such as epidemic outbreaks, financial bubbles, or runaway inflation, *more capable* language models produce *worse* forecasts.
- Under superlinear growth (e.g., exponential doubling early in an epidemic), stronger models (e.g., Llama-3.1 405B) detect the trend earlier than weaker ones (e.g., 70B).
- But their superior pattern recognition leads to aggressive extrapolation: they project the rising curve far into the future with overconfidence, inflating the upper tail of their predicted distributions.
- Linear growth: stronger models remain stable—capability still improves accuracy.
- Exponential growth + trend reversal: strong models' CRPS scores collapse, sometimes performing worse than smaller models from years earlier.
- How does the model internally decide when to stop extrapolating?
- Is the aggressive tendency learned from pretraining data that over-represents "growth mythology," or from post-training (RLHF) rewarding confident, deterministic answers?
Regime change: the logical blind spot
The real world features regime changes—a pandemic can stop abruptly due to lockdowns; a market can crash on sudden policy shifts. When reality brakes or reverses, the strong model's inflated upper tail becomes a massive forecast error.
Experimental results:
The unexplained "extrapolation black box"
The paper diagnoses the problem but leaves the mechanism open:
Takeaway
Capability can become a cognitive curse. The paper undercuts blind faith that bigger models solve everything: in high-stakes decisions, a powerful AI may be the last one to hit the brakes at a sharp turn. When an AI predicts unlimited industry growth or world-ending risk, remember it may simply be extending a line drawn from past experience—extrapolating beyond the edge of what data can justify.
True foresight means seeing past the boom to the ending.