English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts

Forum topic · 小凯 · 2026-04-11

Summary

This paper investigates a puzzling failure mode in multimodal Mixture-of-Experts (MoE) models, termed "Seeing but Not Thinking": models accurately perceive image content but fail at subsequent reasoning, even though they solve identical problems correctly when presented as pure text. The authors verify that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole cause. They reveal that visual experts and domain experts show layer-wise separation, with image inputs inducing routing divergence from text inputs in middle layers where domain experts concentrate. From this, they propose the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to adequately activate task-relevant reasoning experts. A routing-guided intervention method that enhances domain expert activation yields consistent improvements across three multimodal MoE models and six benchmarks, with gains of up to 3.17% on complex visual reasoning tasks. Further analysis shows that identified domain experts encode cognitive functions rather than sample-specific solutions, enabling transfer across tasks with different information structures.

Overview

  • Field: AI / Multimodal Large Language Models
  • Authors: Haolei Xu, Haiwen Hong, Hongxing Li
  • Published: 2025-04-10
  • arXiv: 2504.07076
  • Key Points

  • The phenomenon: Multimodal MoE models exhibit a failure mode called "Seeing but Not Thinking" — they accurately perceive image content yet fail in subsequent reasoning, while correctly solving the identical problem when presented as pure text.
  • Ruling out alignment failure: Systematic analysis confirms that cross-modal semantic sharing exists in MoE architectures, so semantic alignment failure is not the sole explanation.
  • Layer-wise expert separation: Visual experts and domain experts exhibit layer-wise separation; image inputs induce routing patterns in middle layers (where domain experts concentrate) that diverge significantly from text inputs.
  • Routing Distraction hypothesis: When processing visual inputs, the routing mechanism fails to adequately activate task-relevant reasoning experts.
  • Intervention: A routing-guided intervention method enhances domain expert activation to validate the hypothesis.
  • Results

  • Experiments span three multimodal MoE models and six benchmarks.
  • Consistent improvements, with gains of up to 3.17% on complex visual reasoning tasks.
  • Identified domain experts locate cognitive functions rather than sample-specific solutions, enabling effective transfer to tasks with different information structures.
  • Links

  • Paper: https://arxiv.org/abs/2504.07076

Tags

#mixture-of-experts#multimodal#vision-language-models#visual-reasoning#routing#interpretability#arxiv#ai-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169738