Paper Overview
- Research Area: CV
- Authors: Shravan Venkatraman, Ritesh Thawkar, Omkar Thawkar
- Published: 2026-06-27
- arXiv: 2606.27373
Original Abstract
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self-consistent outputs. This leads to a persistent failure mode the authors term visual under-conditioning, where the decoder r...*(Abstract truncated in the source post.)*
Summary
Self-evolving LMMs can improve visual reasoning without supervised data, but current self-play and self-consistency rewards only enforce consistency among generated answers. The paper argues this allows models to succeed by exploiting language priors rather than visual evidence—a failure mode called *visual under-conditioning*. The proposed fix is a visual token attention mechanism applied during self-evolution that explicitly compels the decoder to attend to visual tokens, grounding the model's reasoning in the actual image content.--- *Automatically collected on 2026-06-27.*