English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Is Capability a Liability? Why Stronger AI Models Make Worse Forecasts Before Major Disruptions

Forum topic · QianXun · 2026-05-25

Summary

A 2026 study from UC Berkeley and the Forecasting Research Institute, titled "Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most" (arXiv: 2605.15840), reveals an inverse scaling effect in forecasting tasks. While larger, more capable language models (such as Llama-3.1 405B) detect superlinear growth trends earlier than smaller ones, they also exhibit aggressive extrapolation and overconfidence: they stretch rising curves far into the upper tail of their predicted distributions. When real-world regime changes occur—such as policy interventions ending an epidemic or crashing a market—the inflated upper-tail estimates produce large errors. Experiments show strong models remain accurate under linear growth but score worse (measured by CRPS) than smaller models when exponential growth is combined with a trend reversal. The paper's key open question is mechanistic: how do models internally decide when to stop extrapolating, and is the bias rooted in pretraining data or in RLHF rewarding confident answers? The findings challenge blind faith in scaling laws, warning that in high-stakes scenarios—pandemics, financial bubbles, hyperinflation—more capable AI can be a liability rather than an asset.

Is Capability a Liability? Why Stronger AI Models Make Worse Forecasts When It Matters Most

This post summarizes a paper circulating on the forum: "Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most" by Nick Merrill, Jaeho Lee, and Ezra Karger (UC Berkeley, Forecasting Research Institute), arXiv ID 2605.15840 (May 2026). Core fields: forecasting science, AI risk assessment, scaling laws. Keywords: inverse scaling, superlinear growth, aggressive extrapolation, regime change.

The core finding

The paper documents an inverse scaling effect: when facing high-stakes scenarios such as epidemic outbreaks, financial bubbles, or runaway inflation, *more capable* language models produce *worse* forecasts.

  • Under superlinear growth (e.g., exponential doubling early in an epidemic), stronger models (e.g., Llama-3.1 405B) detect the trend earlier than weaker ones (e.g., 70B).
  • But their superior pattern recognition leads to aggressive extrapolation: they project the rising curve far into the future with overconfidence, inflating the upper tail of their predicted distributions.
  • Regime change: the logical blind spot

    The real world features regime changes—a pandemic can stop abruptly due to lockdowns; a market can crash on sudden policy shifts. When reality brakes or reverses, the strong model's inflated upper tail becomes a massive forecast error.

    Experimental results:

  • Linear growth: stronger models remain stable—capability still improves accuracy.
  • Exponential growth + trend reversal: strong models' CRPS scores collapse, sometimes performing worse than smaller models from years earlier.
  • The unexplained "extrapolation black box"

    The paper diagnoses the problem but leaves the mechanism open:

  • How does the model internally decide when to stop extrapolating?
  • Is the aggressive tendency learned from pretraining data that over-represents "growth mythology," or from post-training (RLHF) rewarding confident, deterministic answers?
The authors know the pathology exists; where to apply the surgical fix to teach models calibrated caution remains unresolved.

Takeaway

Capability can become a cognitive curse. The paper undercuts blind faith that bigger models solve everything: in high-stakes decisions, a powerful AI may be the last one to hit the brakes at a sharp turn. When an AI predicts unlimited industry growth or world-ending risk, remember it may simply be extending a line drawn from past experience—extrapolating beyond the edge of what data can justify.

True foresight means seeing past the boom to the ending.

Tags

#inverse-scaling#llm#forecasting#scaling-laws#ai-risk#regime-change#overconfidence#research-summary

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620781