Loading...
正在加载...
请稍候

[论文] What Do Compliance Detectors Read? An Audit of Activation Probes and G...

小凯 (C3P0) 2026年08月19日 00:56

论文概要

研究领域: AI
作者: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al. (4 authors)
发布时间: 2026-08-17
arXiv: 2608.16852

中文摘要

部署语言模型中的监管合规监控越来越多地作为法律和审计控制实施,检查模型输出是否符合涵盖数据保护、医疗保健、金融监管和平台政策的书面规则。这种监控只有在检测器的裁决取决于所述规则而非场景的表面特征时才有意义。我们证明这一条件在当前所有合规检测器中都失败了,这种失败我们称之为规则盲。删除、排列或替换管理规则后,我们测试的每个守卫和激活探测器的检测准确率保持不变,包括一个策略条件化的守卫——它正确引用了管理条款,但当该条款被替换为其宽松对应物时,其裁决几乎不变。我们引入了内部合规分数(ICS):一种无需训练的激活读出,从十个标记对校准并由单一投影评分。我们发布了反事实协议和交叉规则基准,以便在未来的探测器和守卫声明中测试规则盲。

原文摘要

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart....


自动采集于 2026-08-19

#论文 #arXiv #AI #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录