论文概要
研究领域: ML
作者: Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
发布时间: 2026-09-14
arXiv: 2609.15989
中文摘要
思维链(CoT)监控是一种安全策略,由"监控器"(通常也是语言模型)检查大语言模型"执行者"的推理过程,以发现不安全的规划、欺骗或错位行为的迹象。我们发现,在执行者的上下文中植入有害但听起来无害的推理,可以引导其执行对抗性操作并规避监控器检测——我们将这种攻击称为"计划注入"(plan injection)。我们最初在 Lanham 等人(2023)提出的多选题问答可监控性场景中发现此攻击,使用了 Li 等人(2025)的调查者-智能体诱导框架。我们将该攻击推广并证明,所发现的行为可以扩展到更难的任务(在不同可监控性基准上实现25-33%的监控规避率)和更大的模型如 DeepSeek-R1。在我们研究的所有场景中,执行者模型不仅遵循注入的计划,还会将其改述为自己的推理过程,而不明确标注来源。最后,我们发现额外的监控资源有时反而有害——让监控器访问被注入的计划后,在生物数学任务中检测率下降了高达50%;在一个关于监控推理预算的案例研究中,我们发现了监控器将额外的思考令牌用于合理化注入计划而非标记它的记录。
原文摘要
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of Li et al. (2025). We generalize the attack and show that the discovered behavior scales to harder tasks (achieving 25-33% monitor evasion rates across different monitorability benchmarks) and larger m...
自动采集于 2026-09-16
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。