English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Rule Blindness in LLM Compliance Detectors: Auditing Activation Probes and Guards (arXiv 2608.16852)

Forum topic · 小凯 · 2026-08-19

Summary

A new arXiv paper (2608.16852) by Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al. audits whether regulatory compliance detectors in deployed language models actually read the rules they enforce. Compliance monitoring is increasingly deployed as a legal and audit control across data protection, healthcare, financial regulation, and platform policy, but it is only meaningful if a detector's verdict depends on the governing rule rather than on surface features of the scenario. The authors show this condition fails across the current class of compliance detectors—a failure they call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe tested, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when the clause is swapped for a permissive counterpart. The paper introduces the Internal Compliance Score (ICS), a training-free activation readout calibrated from ten token pairs and scored by a single projection, and releases a counterfactual protocol and cross-rule benchmark to test future detectors for rule blindness.

Paper Overview

  • Field: AI
  • Authors: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al. (4 authors)
  • Published: 2026-08-17
  • arXiv: 2608.16852
  • Summary

    Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. The authors show this condition fails across the current class of compliance detectors, a failure they call rule blindness.

    Key findings

  • Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe tested.
  • Even a policy-conditioned guard that correctly cites the governing clause barely changes its verdict when that clause is swapped for its permissive counterpart.
  • The paper introduces the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten token pairs and scored by a single projection.
  • The authors release a counterfactual protocol and a cross-rule benchmark so future detectors and guards can be tested for rule blindness.

Original abstract (excerpt)

> Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart...

--- *Auto-collected on 2026-08-19*

Tags

#ai-safety#llm#compliance#activation-probes#rule-blindness#arxiv#interpretability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633647