English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts Models

Forum topic · 小凯 · 2026-04-12

Summary

This paper investigates a puzzling failure mode in multimodal Mixture-of-Experts (MoE) models called Seeing but Not Thinking: models accurately perceive image content yet fail at subsequent reasoning, even though they solve the identical problem correctly when presented as pure text. Through systematic analysis, the authors first verify that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole explanation. They then reveal that visual experts and domain experts exhibit layer-wise separation, with image inputs inducing significant routing divergence from text inputs in middle layers where domain experts concentrate. Based on these findings, the paper proposes the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to sufficiently activate task-relevant reasoning experts. To validate this, the authors design a routing-guided intervention that enhances domain expert activation. Experiments across three multimodal MoE models and six benchmarks show consistent improvements, with gains up to 3.17% on complex visual reasoning tasks. Further analysis shows that domain expert identification captures cognitive functions rather than sample-specific solutions, enabling effective transfer across tasks with different information structures. arXiv: 2504.07859.

Paper Overview

Field: NLP Authors: Haolei Xu, Haiwen Hong, Hongxing Li Published: 2025-04-10 arXiv: 2504.07859

Abstract

Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, the authors identify a puzzling phenomenon termed Seeing but Not Thinking: models accurately perceive image content yet fail in subsequent reasoning, while correctly solving identical problems presented as pure text.

Through systematic analysis, the paper first verifies that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole explanation. It then reveals that visual experts and domain experts exhibit layer-wise separation, with image inputs inducing significant routing divergence from text inputs in middle layers where domain experts concentrate.

Based on these findings, the authors propose the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to sufficiently activate task-relevant reasoning experts. To validate this hypothesis, they design a routing-guided intervention to enhance domain expert activation.

Key Results

  • Experiments across three multimodal MoE models and six benchmarks show consistent improvements, with gains of up to 3.17% on complex visual reasoning tasks.
  • Analysis reveals that domain expert identification localizes cognitive functions rather than sample-specific solutions, enabling effective transfer across tasks with different information structures.
  • Links

  • arXiv paper: https://arxiv.org/abs/2504.07859

Tags

#mixture-of-experts#multimodal-models#vision-language#reasoning#interpretability#arxiv#nlp#routing-distraction

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169764