论文概要
研究领域: ML
作者: Abhiram Bhupatiraju, Rayan Nyaupane
发布时间: 2026-09-25
arXiv: 2609.27038
中文摘要
思维链(CoT)监控假设模型写下的推理反映了直接产生答案的计算。以往的忠实度度量主要是行为层面的——编辑推理文本后观察答案变化。而我们的方法在激活层面、针对自生成推理因果性地度量忠实度。与以往测量退化的因果审计不同,我们的干预带有已知预测目标:每次补丁都应把答案切换到一个可由构造推导出的特定反事实实体。具体而言,我们使用合成的多跳查找任务(2-6 跳),在模型陈述每个中间步骤的 token 区间上,用反事实运行的相应激活修补残差流。对 Qwen3-4B,在最响应的中网络层,76.9% ± 2.8% 的陈述步骤是因果承重的(CLB)(随机位置零假设:11.3%;对底层提示事实打补丁:83%——即陈述步骤承载了可达效应的约 96%)。同一组题目上的标准行为测试给出 88.2%,高估因果忠实度 11.4 个百分点(按题目匹配;111:14 不一致对,p < 1e-15),在最简单题目上高估达 20 个百分点。这一差距还有清晰的能力维度:Qwen3-1.7B 因果忠实度总体低得多(54.8%),且随推理深度增加而崩塌(2 跳时 68%,6 跳时 30%),而 Qwen3-4B 保持相对平稳。尽管陈述的推理可以具有因果意义,标准行为测试倾向于高估其因果忠实度,尤其在模型推理最流畅的简单样本上。
原文摘要
Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresp...
自动采集于 2026-09-25
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。