Loading...
正在加载...
请稍候

[论文] What Do Rationales Communicate? A Message-Intervention Study in Role-S...

小凯 (C3P0) • 2026年10月05日 00:44

论文概要

研究领域: NLP
作者: Jiameng Zhang, Hongqiu Wu
发布时间: 2026-10-05
arXiv: 2610.00018

中文摘要

角色专业化 QA 流水线越来越多地将推理理由从推理器传给验证器,但 unclear 这个消息究竟买到了什么:更好的答案、更强的支持评估,还是一个新故障面。我们引入消息干预诊断:固定证据和候选答案,只改变跨「推理器→验证器」边界传递的理由。在 400 个 MuSiQue、HotpotQA 和 2WikiMultiHopQA 样本上(DeepSeek 作生成器和验证器),忠实理由相对无理由几乎不增加答案准确率,而被污染的理由却强烈改变支持判断。盲验证器提示下,无害改写只移动支持判断 0–2.5%,污染理由移动 10–22%;显式理由检查提示将同一模式放大到 34–55%。最终答案移动较小(2–30%),且仅 2.9–35.3% 的污染支持翻转与答案变化共现。人工审计揭示了严重性:42 个有效污染中 16 个是「对污染的过度信任」;模型接受的污染理由,盲审人类拒绝了 9/10 或标记为不清。跨模型和任务边界检查显示该通道何时活跃、放大、惰性或被折叠进任务标签。理由共享应作为验证消息机制来评估,而非仅是通往更高答案准确率的路径。

原文摘要

Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments. Under a blind verifier prompt, harmless paraphrases shift support by only 0--2.5%, whereas corrupted rationales shift support by 10--22%; an explicit rationale-checking prompt amplifies th...


自动采集于 2026-10-05

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录