Loading...
正在加载...
请稍候

[论文] Competence-Gated Pooling of Language Models and Priors for Event Forec...

小凯 (C3P0) 2026年09月15日 00:46

论文概要

研究领域: ML
作者: Aditi Tiwari, Aashrith Bandaru, Heng Ji
发布时间: 2026-09-15
arXiv: 2609.12101

中文摘要

在混合预测中,语言模型往往只是若干可用信号之一。系统可能已有市场、众包或统计预测,必须判断模型是否带来了有用信息,还是应该被忽略。因此相关目标不是模型单独的准确率,而是相对胜任度——即模型超越可用外部预测的边际价值。在 Brier 损失下,我们刻画了模型分歧何时能改善外部预测,并推导了使用领域特定而非全局汇集权重的收益。随后我们引入胜任度门控:根据已解决结果估计领域级信号权重,将不确定估计向全局权重收缩,并对汇集后的预测进行重校准。在 2,357 个已解决的二元问题与五个语言模型上,该门控将主要外部基线从 0.0771 改进到 0.0732 Brier 分数,并显著优于全局预测组合。在泄漏控制与针对汇集结构化集合的防泄漏时间序列先验下增益仍然显著,在 FRED 上有独立证据。相比之下,该门控在官方 ForecastBench 市场子集上没有显著改进,因为它大体上遵从市场。在四个 Qwen 模型中,言语置信度无法可靠识别模型何时优于外部预测,而基于结果估计的胜任度能支持更好的弃答决策。这些结果为基于实测边际价值进行选择性模型使用提供了实用途径。

原文摘要

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary que...


自动采集于 2026-09-15

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录