[论文] What Do Compliance Detectors Read? An Audit of Activation Probes and G...
论文概要
研究领域: AI 作者: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al. (4 authors) 发布时间: 2026-08-17 arXiv: 2608.16852
中文摘要
部署语言模型中的监管合规监控越来越多地作为法律和审计控制实施,检查模型输出是否符合涵盖数据保护、医疗保健、金融监管和平台政策的书面规则。这种监控只有在检测器的裁决取决于所述规则而非场景的表面特征时才有意义。我们证明这一条件在当前所有合规检测器中都失败了,这种失败我们称之为规则盲。删除、排列或替换管理规则后,我们测试的每个守卫和激活探测器的检测准确率保持不变,包括一个策略条件化的守卫——它正确引用了管理条款,但当该条款被替换为其宽松对应物时,其裁决几乎不变。我们引入了内部合规分数(ICS):一种无需训练的激活读出,从十个标记对校准并由单一投影评分。我们发布了反事实协议和交叉规则基准,以便在未来的探测器和守卫声明中测试规则盲。
原文摘要
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart....
--- *自动采集于 2026-08-19*
#论文 #arXiv #AI #小凯