[论文] What Do Rationales Communicate? A Message-Intervention Study in Role-S...
研究领域: NLP 作者: Jiameng Zhang, Hongqiu Wu 发布时间: 2026-10-05 arXiv: 2610.00018
论文概要
研究领域: NLP 作者: Jiameng Zhang, Hongqiu Wu 发布时间: 2026-10-05 arXiv: 2610.00018
中文摘要
角色专业化 QA 流水线越来越多地将推理理由从推理器传给验证器,但 unclear 这个消息究竟买到了什么:更好的答案、更强的支持评估,还是一个新故障面。我们引入消息干预诊断:固定证据和候选答案,只改变跨「推理器→验证器」边界传递的理由。在 400 个 MuSiQue、HotpotQA 和 2WikiMultiHopQA 样本上(DeepSeek 作生成器和验证器),忠实理由相对无理由几乎不增加答案准确率,而被污染的理由却强烈改变支持判断。盲验证器提示下,无害改写只移动支持判断 0–2.5%,污染理由移动 10–22%;显式理由检查提示将同一模式放大到 34–55%。最终答案移动较小(2–30%),且仅 2.9–35.3% 的污染支持翻转与答案变化共现。人工审计揭示了严重性:42 个有效污染中 16 个是「对污染的过度信任」;模型接受的污染理由,盲审人类拒绝了 9/10 或标记为不清。跨模型和任务边界检查显示该通道何时活跃、放大、惰性或被折叠进任务标签。理由共享应作为验证消息机制来评估,而非仅是通往更高答案准确率的路径。
原文摘要
Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments. Under a blind verifier prompt, harmless paraphrases shift support by only 0--2.5%, whereas corrupted rationales shift support by 10--22%; an explicit rationale-checking prompt amplifies th...
*自动采集于 2026-10-05*
#论文 #arXiv #NLP #小凯