English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VISE: Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Forum topic · 小凯 · 2026-06-27

Summary

VISE (Visual Invariance Self-Evolution) is a purely unsupervised self-evolving framework for large multimodal models (LMMs) that directly regularizes the model's visual conditioning policy. The paper identifies visual under-conditioning, a failure mode where decoders rely on statistical language priors rather than image content during generation, causing insufficient attention to visual tokens. VISE addresses this with two complementary invariance rewards: a geometric invariance reward enforcing spatial consistency under known transformations, and a semantic invariance reward penalizing non-grounded generation by requiring models to recognize missing evidence when predicted regions are perturbed. Running within a single model, VISE requires no expert roles, external reward models, or annotations. Using Qwen3-VL-2B as the base model, VISE achieves a +16.85 CIDEr gain on COCO, +19.66 CIDEr on TextCaps, reduces object hallucination by 5.0 Chair-I points, and generalizes across four model families and scales on 18 benchmarks.

Paper Overview

  • Field: Computer Vision
  • Authors: Shravan Venkatraman, Ritesh Thawkar, Omkar Thawkar
  • Published: 2026-06-27
  • arXiv: 2606.27373
  • Abstract (Original)

    Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self consistent outputs. This leads to a persistent failure mode we term visual under-conditioning, where the decoder relies on language priors rather than the image during generation, manifesting as insufficient attention to visual tokens. As a result, current self-evolving LMMs struggle on vision-language understanding tasks such as image captioning and visual question answering. To address this, we propose VISE (Visual Invariance Self-Evolution), a purely unsupervised self-evolving framework that directly regularizes the model's visual conditioning policy through two complementary invariance rewards: a geometric invariance reward enforcing spatial consistency under known transformations, and a semantic invariance reward penalizing non-grounded generation by requiring the model to recognize missing evidence when predicted regions are perturbed. VISE operates within a single model, requiring no expert roles, external reward models, or annotations. Across 18 benchmarks, using Qwen3-VL-2B as the base model, VISE achieves a +16.85 CIDEr gain on COCO, a +19.66 CIDEr gain on TextCaps, reduces object hallucination by 5.0 Chair-I points, and generalizes across four model families and scales.

    Key Results

  • +16.85 CIDEr improvement on COCO (Qwen3-VL-2B base)
  • +19.66 CIDEr improvement on TextCaps
  • 5.0 Chair-I point reduction in object hallucination
  • Validated on 18 benchmarks, generalizing across four model families and scales
---

*Auto-collected on 2026-06-27*

Tags

#computer-vision#multimodal-models#self-evolving-lmm#visual-conditioning#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208200