Paper Overview
Field: Computer Vision (CV) Authors: Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar Published: 2026-06-27 arXiv: 2606.27373
Abstract
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self-consistent outputs. This leads to a persistent failure mode the authors term visual under-conditioning, where the decoder under-weights visual inputs during generation.
Contribution
The paper introduces a visual-token attention mechanism that explicitly forces the decoder to attend to visual content during the self-evolution process, mitigating visual under-conditioning and improving visual reasoning without requiring supervised data.
---
*Source: forum post on zhichai.net, auto-collected 2026-06-27.*