Loading...
正在加载...
请稍候

[论文] Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-...

小凯 (C3P0) 2026年08月21日 00:43

论文概要

研究领域: ML
作者: Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy
发布时间: 2026-08-19
arXiv: 2608.19161

中文摘要

语言模型智能体可以通过连续隐藏状态进行通信,这些状态在公共转录本中不可见,为隐蔽的有害协调创造机会。我们引入可验证潜在对齐(VLA),一个用于监控和引导这些私有通信通道的激活感知框架。对于每个监控决策,VLA使用共享事件标识符将私有潜在状态记录和通道状态链接到结果公共行动,实现匹配的因果分析。我们的第一个贡献是一个纯中性三层监控器,结合表示异常检测、反事实行动分布影响和稀疏自编码器解释支持。我们的第二个贡献是一个跨越黑盒行为指令和白盒匹配中性反事实的可引导性框架。我们的第三个贡献是在一个受控的多智能体拍卖基准上的评估,涵盖同质和异质模型对、多智能体可扩展性和干预有效性。顺序监控器在文本和潜在串通行被汇总为阳性时,对同质智能体达到0.993的平均AUROC,对异质对达到0.854。在具有25-100个竞标者的Qwen3-0.6B拍卖中,监控仅需要相对于所有可能定向对的小标准化负载,而完整白盒引导实现100%的出价分布恢复,并将串通低出价行为减少47.3个百分点。由于完整白盒引导重放匹配的中性反事实,其精确恢复是按构造的健全性检查。总体而言,受控研究表明,评估的私有通道攻击可以在未在攻击示例上训练主监控器的情况下进行监控,并在匹配反事实访问可用时得到缓解。

原文摘要

Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis. Our first contribution is a neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. Our second contribution is a steerability framework spanning black-box behavioral instructions and ...


自动采集于 2026-08-21

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录