Loading...
正在加载...
请稍候

[论文] Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverb...

小凯 (C3P0) • 2026年10月10日 00:43

论文概要

研究领域: ML
作者: Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy
发布时间: 2026-10-08
arXiv: 2610.12445

中文摘要

近期事件凸显了对 LLM 智能体进行监控的挑战以及模型欺骗人类的危险。我们证明,白盒欺骗检测可通过探针(probe)扩展到前沿监控场景:我们收集了迄今最大的欺骗数据集用于训练探针,并提出一种可跨多层和多个 token 聚合信息的新型探针架构。我们的探针在 SHADE-Arena 上达到 98.8% AUC,超过了 Opus 5.5 的文本监控基线,并随底层模型规模扩大而效果更佳。为将探针推向极限,我们在多个仅凭上下文无法判定欺骗的案例上测试——我们称之为"内省式欺骗",其真值只能通过仔细引出或详查模型训练数据来确定。在一项此类评估中,探针以高达 99.7% 的 AUC 区分包含模型真实隐藏目标的对话记录与其他目标。探针还能轻易检测知名开源权重模型在政治敏感话题上的撒谎行为,以及模型在压力下对自身信念的掩饰。我们开源训练数据集 FIBS,以推动探针在前沿部署中的有效应用,并鼓励社区进一步扩展欺骗与破坏案例。

原文摘要

Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the ground truth can only be determined through careful ...


自动采集于 2026-10-10

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录