Loading...
正在加载...
请稍候

[论文] Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confi...

小凯 (C3P0) 2026年09月22日 00:45

论文概要

研究领域: NLP
作者: Andre Bacellar
发布时间: 2026-09-18
arXiv: 2609.22056

中文摘要

多跳检索的失败并非在查询间均匀分布,而是聚集在结构可预测的子群体中。我们证明了两条结果来形式化这一结构。第一(CWAR 可归约性):当且仅当检索特征携带关于成功与否的互信息时,自信失败的降低才是可实现的;LLM-judge 流水线满足这一条件,但在纯 dense 设置中明显更弱,这解释了两种机制之间的 AUC-AC 差距。第二(特征机制互补性):没有任何单一 ANN 分数特征能在所有失败机制中取得最佳预测表现;主导特征因数据集而异(MuSiQue 上是查询长度,HoVer 上是 hop-1 集中度),且一个构造性见证对表明每个特征在一种机制中是必要的,在另一种中则无贡献。我们将这些原则实例化为 RegimeAbstain:它计算检索置信度分数(RCS)——一个由至多九个查询-ANN 结构特征构成的 logistic 函数,全部特征无需任何额外 LLM 调用即可获得——并用其实施经校准的弃权策略。我们定义了自信错误回答率(CWAR)指标,在三个多跳基准(MuSiQue、2WikiMultiHopQA、HoVer)和两种检索架构(LLM-judge 与纯 dense)上评估,覆盖 CWAR 从 14.5% 到 62.1% 的五种失败机制。RCS 在全部五种条件下均取得最佳或并列最佳的 AUC-AC(对比八个置信度基线)。在 MuSiQue(LLM-judge)上,RCS 在 50% 覆盖率下将 CWAR 从 39.5% 降至 20.6%(相对降低 47.8%),ECE=0.035。在 MuSiQue 上训练的模型迁移到 2WikiMultiHopQA 时 AUC 仅损失 0.5 个百分点,证实了机制特征具有跨领域结构。

原文摘要

Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the othe...


自动采集于 2026-09-22

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录