Paper Overview
Field: NLP Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar Published: 2026-08-11 arXiv: 2508.03805
Summary
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control.
The authors introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses:
1. Feature isolation — diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training 2. Task-specific feature detection — per-token contrastive triggering analysis to isolate causal features 3. Feature-level control — causal ablation or steering of discovered feature directions
Results
The authors trained multimodal SAEs for three MLLM families — LLaVA-MORE, PaliGemma 2, and InternVL3.5 — and evaluated on visual spatial understanding, multimodal safety, and OCR:
- Ablation of discovered sparse, task-specific features selectively degraded targeted behavior: -12% on spatial tasks, -17% on OCR, and -24% attack success rate on multimodal safety attacks, without affecting VQA performance
- Steering these features improved spatial and OCR accuracy by an average of +3.6% and +1.8% compared to standard single-layer steering baselines
Conclusion
These results demonstrate that multimodal SAEs can serve not only as interpretability tools but also as mechanisms for auditing, steering, and controlling MLLM behavior toward safer and more capable generation.
--- *Auto-collected on 2026-08-12*