English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MMDiff: Multimodal Model Diffing for Feature Discovery and Control in MLLMs

Forum topic · 小凯 · 2026-08-11

Summary

MMDiff is a multimodal model-diffing framework that trains sparse autoencoders (SAEs) on multimodal large language models (MLLMs) and turns them into feature-level interfaces for discovering and controlling multimodal behavior. Proposed by Hunar Batra, Lachin Naghashyar, and Ashkan Khakzar (arXiv:2508.03805), the framework supports three uses: (i) feature isolation by diffing a base-LM SAE against its multimodal-adapted counterpart, (ii) task-specific feature detection via per-token contrastive trigger analysis, and (iii) feature-level control through causal ablation or steering of discovered feature directions. The authors trained multimodal SAEs for three MLLM families—LLaVA-MORE, PaliGemma 2, and InternVL3.5—and evaluated on visual spatial understanding, multimodal safety, and OCR. Ablating discovered features selectively degraded targeted behavior by 12% on spatial tasks, 17% on OCR, and reduced attack success rate by 24% on multimodal safety attacks without affecting VQA performance. Steering these features improved spatial and OCR accuracy by +3.6% and +1.8% on average over standard single-layer steering baselines. Results show multimodal SAEs can serve not only as interpretability tools but also as mechanisms for auditing, steering, and controlling MLLM behavior for safer, more capable generation.

Paper Overview

Field: NLP Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar Published: 2026-08-11 arXiv: 2508.03805

Summary

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control.

The authors introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses:

1. Feature isolation — diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; 2. Task-specific feature detection — per-token contrastive trigger analysis to isolate causal features; 3. Feature-level control — causal removal or steering of discovered feature directions.

Key Results

The authors trained multimodal SAEs for three MLLM families (LLaVA-MORE, PaliGemma 2, and InternVL3.5) and evaluated on visual spatial understanding, multimodal safety, and OCR:

  • MMDiff discovers sparse, task-specific features whose removal selectively degrades target behavior by 12% on spatial tasks, 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks — all without affecting VQA performance.
  • Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over standard single-layer steering baselines.

Conclusion

These results show that multimodal SAEs can serve not only as interpretability tools but also as mechanisms for auditing, steering, and controlling MLLM behavior for safer, more capable generation.

--- *Auto-collected on 2026-08-12.*

Tags

#mllm#sparse-autoencoders#interpretability#model-diffing#multimodal#ai-safety#steering#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633348