[论文] Competence-Gated Pooling of Language Models and Priors for Event Forec...
研究领域: ML 作者: Aditi Tiwari, Aashrith Bandaru, Heng Ji 发布时间: 2026-09-15 arXiv: 2609.12101
论文概要
研究领域: ML 作者: Aditi Tiwari, Aashrith Bandaru, Heng Ji 发布时间: 2026-09-15 arXiv: 2609.12101
中文摘要
在混合预测中,语言模型往往只是若干可用信号之一。系统可能已有市场、众包或统计预测,必须判断模型是否带来了有用信息,还是应该被忽略。因此相关目标不是模型单独的准确率,而是相对胜任度——即模型超越可用外部预测的边际价值。在 Brier 损失下,我们刻画了模型分歧何时能改善外部预测,并推导了使用领域特定而非全局汇集权重的收益。随后我们引入胜任度门控:根据已解决结果估计领域级信号权重,将不确定估计向全局权重收缩,并对汇集后的预测进行重校准。在 2,357 个已解决的二元问题与五个语言模型上,该门控将主要外部基线从 0.0771 改进到 0.0732 Brier 分数,并显著优于全局预测组合。在泄漏控制与针对汇集结构化集合的防泄漏时间序列先验下增益仍然显著,在 FRED 上有独立证据。相比之下,该门控在官方 ForecastBench 市场子集上没有显著改进,因为它大体上遵从市场。在四个 Qwen 模型中,言语置信度无法可靠识别模型何时优于外部预测,而基于结果估计的胜任度能支持更好的弃答决策。这些结果为基于实测边际价值进行选择性模型使用提供了实用途径。
原文摘要
In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary que...
*自动采集于 2026-09-15*
#论文 #arXiv #ML #小凯