Paper Overview
Research Area: Computer Vision (CV) Authors: Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar Published: 2026-06-27 arXiv: 2606.27373
Abstract
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self-consistent outputs. This leads to a persistent failure mode the authors term visual under-conditioning, where the decoder produces plausible-looking answers that are not actually grounded in the visual input.
Proposed Approach
The paper introduces a visual token attention mechanism that explicitly enforces the decoder to attend to visual content during the self-evolution process, addressing the visual under-conditioning problem at its source.
---
*Auto-collected on 2026-06-27.*