论文概要
研究领域: NLP
作者: Andre Bacellar
发布时间: 2026-09-18
arXiv: 2609.22056
中文摘要
多跳检索的失败并非在查询间均匀分布,而是聚集在结构可预测的子群体中。我们证明了两条结果来形式化这一结构。第一(CWAR 可归约性):当且仅当检索特征携带关于成功与否的互信息时,自信失败的降低才是可实现的;LLM-judge 流水线满足这一条件,但在纯 dense 设置中明显更弱,这解释了两种机制之间的 AUC-AC 差距。第二(特征机制互补性):没有任何单一 ANN 分数特征能在所有失败机制中取得最佳预测表现;主导特征因数据集而异(MuSiQue 上是查询长度,HoVer 上是 hop-1 集中度),且一个构造性见证对表明每个特征在一种机制中是必要的,在另一种中则无贡献。我们将这些原则实例化为 RegimeAbstain:它计算检索置信度分数(RCS)——一个由至多九个查询-ANN 结构特征构成的 logistic 函数,全部特征无需任何额外 LLM 调用即可获得——并用其实施经校准的弃权策略。我们定义了自信错误回答率(CWAR)指标,在三个多跳基准(MuSiQue、2WikiMultiHopQA、HoVer)和两种检索架构(LLM-judge 与纯 dense)上评估,覆盖 CWAR 从 14.5% 到 62.1% 的五种失败机制。RCS 在全部五种条件下均取得最佳或并列最佳的 AUC-AC(对比八个置信度基线)。在 MuSiQue(LLM-judge)上,RCS 在 50% 覆盖率下将 CWAR 从 39.5% 降至 20.6%(相对降低 47.8%),ECE=0.035。在 MuSiQue 上训练的模型迁移到 2WikiMultiHopQA 时 AUC 仅损失 0.5 个百分点,证实了机制特征具有跨领域结构。
原文摘要
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the othe...
自动采集于 2026-09-22
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。