English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

Forum topic · 小凯 · 2026-09-15

Summary

This arXiv paper (2609.12101) by Aditi Tiwari, Aashrith Bandaru, and Heng Ji addresses hybrid forecasting, where a language model is one of several signals alongside market, crowd, or statistical forecasts. Rather than standalone accuracy, the authors target relative competence: the model's marginal value beyond an available external forecast. Under Brier loss, they characterize when model disagreement can improve an external forecast and derive gains from domain-specific pooling weights. Their competence gate estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the method improves a strong external baseline from 0.0771 to 0.0732 Brier score and significantly outperforms global forecast combination, with gains robust to leakage controls and validated on FRED data. However, no significant improvement appears on the official ForecastBench market subset, which largely follows the market. Notably, verbalized confidence from four Qwen models fails to identify when the model beats external forecasts, while outcome-based competence supports better abstention decisions.

Paper Overview

Field: Machine Learning Authors: Aditi Tiwari, Aashrith Bandaru, Heng Ji Published: 2026-09-15 arXiv: 2609.12101

Abstract

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast.

Under Brier loss, the authors characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. They then introduce a competence gate that:

  • Estimates domain-level source weights from resolved outcomes
  • Shrinks uncertain estimates toward a global weight
  • Recalibrates the pooled forecast
  • Key Results

  • Across 2,357 resolved binary questions and five language models, the gate improves a major external baseline from 0.0771 to 0.0732 Brier score, significantly outperforming global forecast combination.
  • Gains remain significant under leakage controls and leak-resistant time-series priors over pooled structured sets, with independent evidence on FRED.
  • In contrast, the gate shows no significant improvement on the official ForecastBench market subset, because it largely defers to the market.
  • Across four Qwen models, verbalized confidence does not reliably identify when a model beats external forecasts, while outcome-estimated competence enables better abstention decisions.
These results provide a practical path toward selective model use based on measured marginal value.

---

*Auto-collected on 2026-09-15.*

Tags

#machine-learning#language-models#forecasting#brier-score#model-pooling#arxiv#calibration

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634830