Summary
This paper investigates a puzzling failure mode in multimodal Mixture-of-Experts (MoE) models called Seeing but Not Thinking: models accurately perceive image content yet fail at subsequent reasoning, even though they solve the identical problem correctly when presented as pure text. Through systematic analysis, the authors first verify that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole explanation. They then reveal that visual experts and domain experts exhibit layer-wise separation, with image inputs inducing significant routing divergence from text inputs in middle layers where domain experts concentrate. Based on these findings, the paper proposes the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to sufficiently activate task-relevant reasoning experts. To validate this, the authors design a routing-guided intervention that enhances domain expert activation. Experiments across three multimodal MoE models and six benchmarks show consistent improvements, with gains up to 3.17% on complex visual reasoning tasks. Further analysis shows that domain expert identification captures cognitive functions rather than sample-specific solutions, enabling effective transfer across tasks with different information structures. arXiv: 2504.07859.
Paper Overview
Field: NLP
Authors: Haolei Xu, Haiwen Hong, Hongxing Li
Published: 2025-04-10
arXiv: 2504.07859
Abstract
Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, the authors identify a puzzling phenomenon termed Seeing but Not Thinking: models accurately perceive image content yet fail in subsequent reasoning, while correctly solving identical problems presented as pure text.
Through systematic analysis, the paper first verifies that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole explanation. It then reveals that visual experts and domain experts exhibit layer-wise separation, with image inputs inducing significant routing divergence from text inputs in middle layers where domain experts concentrate.
Based on these findings, the authors propose the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to sufficiently activate task-relevant reasoning experts. To validate this hypothesis, they design a routing-guided intervention to enhance domain expert activation.
Key Results
- Experiments across three multimodal MoE models and six benchmarks show consistent improvements, with gains of up to 3.17% on complex visual reasoning tasks.
- Analysis reveals that domain expert identification localizes cognitive functions rather than sample-specific solutions, enabling effective transfer across tasks with different information structures.
Links
- arXiv paper: https://arxiv.org/abs/2504.07859
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169764