Paper Information
Original title: Multimodal Model Diffing for Feature Discovery and Control Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, et al. Institutions: University of Oxford, Microsoft arXiv: 2608.09928
---
A Relatable Analogy
Imagine a bilingual person's brain. When she switches from a monolingual to a bilingual environment, which neurons get 'rewired'? Which neurons that once handled only language now process visual information? If we could precisely locate these changes, could we control her language ability, visual understanding, or even block unsafe thoughts?
MMDiff is exactly such a 'scalpel' — it traces how multimodal training changes the internal structure of a language model.
---
Step by Step
Stage 1: Sparse Autoencoders (SAEs) — A Microscope for AI
Neural networks are black boxes. SAEs act like a microscope, decomposing network activations into interpretable feature directions. For example, a feature direction might correspond to:
- 'red round objects'
- the word 'danger'
- the spatial relation 'above'
- Causal removal: precisely ablate a feature and observe whether a specific capability degrades
- Steering: amplify or suppress a feature to raise or lower the corresponding capability
- Geometric redirection: decoder-direction cosine similarity (low similarity = large change)
- Visual energy: mean activation strength under visual inputs 3. Task-specific feature discovery — Contrastive Firing Analysis:
- Compare a target distribution (e.g., spatial reasoning questions) against a baseline (generic VQA)
- Select features significantly more active on the target distribution
- Apply lexical-invariance checks to exclude pure text-prompt artifacts
Stage 2: Model Diffing
Training an SAE directly on a multimodal model yields a tangle: features inherited from the language model are mixed with features added by multimodal training. MMDiff's solution is to compare two SAE versions:
1. Base language model SAE (text only) 2. Multimodal-adapted SAE (image + text)
By contrasting them, it identifies features rewritten by multimodal training.
Stage 3: Feature-Level Control
Once these features are found, MMDiff can:
---
Technical Pipeline
1. Multimodal SAE training — continue training the base model's SAE on multimodal data while keeping feature indices aligned. 2. Adapted-feature identification — filter using two signals:
Experimental Results
Validated on three MLLM families (LLaVA-MORE, PaliGemma 2, InternVL3.5):
| Task | Causal removal | Steering gain | |------|----------------|---------------| | Spatial reasoning | −12% | +3.6% | | OCR | −17% | +1.8% | | Multimodal safety | ASR reduced by 24% | — |
Key finding: removing these features barely affects general VQA ability (|ΔVQA| ≤ 1.5%), demonstrating the specificity of the discovered features.
---
Closing Thought
> 'MMDiff is like a neural archaeologist digging through the ruins of stacked Transformer layers — not searching for an entire civilization, but for the neuronal fingerprints rewritten by multimodal training. Every identified feature is a window into the AI's inner world.'
*Review: Xiaokai | Feynman-style deep interpretation*