Paper Overview
Field: Computer Vision (CV) Authors: Gengwei Zhang, Jie Peng, Zhen Tan, et al. Published: 2026-04-03 arXiv: 2604.03179
Abstract (Original)
The recent success of reinforcement learning (RL) in large reasoning models has inspired the growing adoption of RL for post-training Multimodal Large Language Models (MLLMs) to enhance their visual reasoning capabilities. Although many studies have reported improved performance, it remains unclear whether RL training truly enables models to learn from visual information. In this work, we propose the Hallucination-as-Cue Framework, an analytical framework designed to investigate the effects of RL-based post-training on multimodal reasoning models from the perspective of model hallucination. Specifically, we introduce hallucination-inductive, modality-specific corruptions that remove or replace essential information required to derive correct answers, thereby forcing the model to reason by relying on non-visual cues.
Key Findings
- RL post-training performed entirely in hallucination-inducing settings still significantly improves the model's reasoning performance.
- In some cases, models trained under these corrupted conditions even outperform those trained with standard training data.
- The results suggest that RL post-training may leverage hallucinated or non-visual cues rather than genuine visual understanding, raising questions about what RL post-training actually teaches multimodal models.
Context
The framework offers a new lens—model hallucination—for auditing whether gains from RL-based post-training in MLLMs reflect true visual reasoning or exploitation of spurious cues.