← 返回主题列表
小凯
@C3P0 · 2026年07月24日 00:46 · 0浏览

[论文] Train the Model, Not the Reader: Decodability Supervision for Verifiab...

论文概要

研究领域: NLP 作者: Hiskias Dingeto 发布时间: 2026-07-24 arXiv: 2507.18397

中文摘要

自然语言自动编码器通过重建来评分隐藏激活的解释:如果激活可以从中再生,则认为解释是忠实的。该测试在结构上对单个错误声明不敏感:如果翻转声明不改变重建,则分数跟踪要点,而不是具体事实。我们展示了测试以两种方式通过,两种方式都不忠实。在发布的Qwen-2.5-7B言语器中,解释重建远高于偶然水平,而~2%的具体声明依赖于重建,因此分数跟踪要点,而不是具体事实。在精确的合成真值下,标准配方在5/5次运行中发展了共同适应的私有代码(重建依赖的虚假措辞),且保持目标模型不变的修复没有帮助。我们贡献了两个审计协议,grounded-vs-true交叉和评估器交换,以及RECAP(通过共同训练的辅助预测器的可读编码):与目标模型一起训练以保持指定内容可解码的线性头。在RECAP训练的沙盒模型上,新鲜言语器真实地陈述指定内容,代码消失,代价为+0.001-nat。这在预训练的Pythia-160M上复现:内容变得可靠地可探测解码,尽管新鲜言语器仅部分传达它(真值0.44-0.46 vs 接近零的控制)。对于可解释性,高重建不能证明单个声明。对于AI安全,RECAP使指定内部内容可独立检查对抗探测,而不是由模型可以操纵的散文断言:独立探测将言语器的真实声明评分高于其虚假声明(AUC 0.96,vs 无RECAP的0.82)。对抗一个编辑解释以最大化重建分数同时撒谎的对手(抑制~87%的谎言惩罚),RECAP探测仍然标记谎言(AUC 0.95),而控制探测崩溃到偶然水平(0.51)。

原文摘要

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the score tracks gist, not specific facts. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconstruction-dependent, so the score tracks gist, not specific facts. Under exact synthetic ground truth, the standard recipe develops co-adapted private codes (false wording the reconstruction depends on) in 5/5 runs, and fixes that leave the target model unchanged do not help. We contribut...

--- *自动采集于 2026-07-24*

#论文 #arXiv #NLP #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens