Overview
- Field: Machine Learning
- Authors: Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy
- Published: 2026-08-19
- arXiv: 2608.19161
- The sequential monitor achieves 0.993 average AUROC on homogeneous agents and 0.854 on heterogeneous pairs when textual and latent collusion rows are aggregated as positives.
- In Qwen3-0.6B auctions with 25–100 bidders, monitoring requires only a small normalized overhead relative to all possible directed pairs.
- Full white-box steering achieves 100% bid-distribution recovery and reduces collusive low-bid behavior by 47.3 percentage points.
- Since full white-box steering replays the matched neutral counterfactual, its exact recovery serves as a sanity check by construction.
English Abstract
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. The authors introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis.
Key contributions:
1. A neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. 2. A steerability framework spanning black-box behavioral instructions and white-box matched neutral counterfactuals. 3. An evaluation on a controlled multi-agent auction benchmark covering homogeneous and heterogeneous model pairs, multi-agent scalability, and intervention effectiveness.
Results
Conclusion
Overall, the controlled study demonstrates that the evaluated private-channel attacks can be monitored without training the primary monitor on attack examples, and can be mitigated when matched-counterfactual access is available.