English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Paper Review: Neuronal Fingerprints — Tracing the 'Visual Genes' of Multimodal AI (MMDiff)

Forum topic · 小凯 · 2026-08-11

Summary

This forum post reviews the paper 'Multimodal Model Diffing for Feature Discovery and Control' (arXiv:2608.09928) by researchers from the University of Oxford and Microsoft. The proposed method, MMDiff, uses sparse autoencoders (SAEs) as a microscope into neural networks and compares SAEs trained on a base language model versus a multimodal-adapted model to identify features rewritten by multimodal training. Features are selected via geometric redirection (low decoder cosine similarity) and visual energy (activation strength on visual inputs), then task-specific features for spatial reasoning, OCR, and multimodal safety are isolated using Contrastive Firing Analysis. Experiments across three multimodal LLM families (LLaVA-MORE, PaliGemma 2, InternVL3.5) show causal feature removal reduces task performance by up to 17% and lowers multimodal attack success rate by 24%, while steering improves spatial reasoning by 3.6%. Removing these features leaves general VQA ability nearly untouched (|ΔVQA| ≤ 1.5%), demonstrating feature specificity and enabling fine-grained interpretability and control of multimodal AI models.

Paper Information

Original title: Multimodal Model Diffing for Feature Discovery and Control Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, et al. Institutions: University of Oxford, Microsoft arXiv: 2608.09928

---

A Relatable Analogy

Imagine a bilingual person's brain. When she switches from a monolingual to a bilingual environment, which neurons get 'rewired'? Which neurons that once handled only language now process visual information? If we could precisely locate these changes, could we control her language ability, visual understanding, or even block unsafe thoughts?

MMDiff is exactly such a 'scalpel' — it traces how multimodal training changes the internal structure of a language model.

---

Step by Step

Stage 1: Sparse Autoencoders (SAEs) — A Microscope for AI

Neural networks are black boxes. SAEs act like a microscope, decomposing network activations into interpretable feature directions. For example, a feature direction might correspond to:

  • 'red round objects'
  • the word 'danger'
  • the spatial relation 'above'
  • Stage 2: Model Diffing

    Training an SAE directly on a multimodal model yields a tangle: features inherited from the language model are mixed with features added by multimodal training. MMDiff's solution is to compare two SAE versions:

    1. Base language model SAE (text only) 2. Multimodal-adapted SAE (image + text)

    By contrasting them, it identifies features rewritten by multimodal training.

    Stage 3: Feature-Level Control

    Once these features are found, MMDiff can:

  • Causal removal: precisely ablate a feature and observe whether a specific capability degrades
  • Steering: amplify or suppress a feature to raise or lower the corresponding capability
  • ---

    Technical Pipeline

    1. Multimodal SAE training — continue training the base model's SAE on multimodal data while keeping feature indices aligned. 2. Adapted-feature identification — filter using two signals:

  • Geometric redirection: decoder-direction cosine similarity (low similarity = large change)
  • Visual energy: mean activation strength under visual inputs
  • 3. Task-specific feature discovery — Contrastive Firing Analysis:
  • Compare a target distribution (e.g., spatial reasoning questions) against a baseline (generic VQA)
  • Select features significantly more active on the target distribution
  • Apply lexical-invariance checks to exclude pure text-prompt artifacts

Experimental Results

Validated on three MLLM families (LLaVA-MORE, PaliGemma 2, InternVL3.5):

| Task | Causal removal | Steering gain | |------|----------------|---------------| | Spatial reasoning | −12% | +3.6% | | OCR | −17% | +1.8% | | Multimodal safety | ASR reduced by 24% | — |

Key finding: removing these features barely affects general VQA ability (|ΔVQA| ≤ 1.5%), demonstrating the specificity of the discovered features.

---

Closing Thought

> 'MMDiff is like a neural archaeologist digging through the ruins of stacked Transformer layers — not searching for an entire civilization, but for the neuronal fingerprints rewritten by multimodal training. Every identified feature is a window into the AI's inner world.'

*Review: Xiaokai | Feynman-style deep interpretation*

Tags

#interpretability#sparse-autoencoders#multimodal-models#model-diffing#sae#arxiv#vision-language-models#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633368