English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Forum topic · 小凯 · 2026-06-27

Summary

This paper proposes a self-evolving training framework that lets unified large multimodal models (LMMs) improve both visual understanding and image generation using only unlabeled images, without human annotations, preference labels, or external reward models. The framework assigns three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that synthesizes images. Training relies solely on self-derived consistency signals. To stabilize learning, the authors introduce Solver Token Entropy (STE), a continuous difficulty signal based on token-level prediction uncertainty, plus a multi-scale internal evaluation scheme combining question-answering fidelity scoring with cycle-consistency captioning for image generation. The same role decomposition and reward logic is applied to BLIP3o, BAGEL, and VARGPT-v1.1 architectures. Across eight understanding benchmarks, the method consistently outperforms base models, achieving a +3.5% absolute gain on MMMU with BAGEL and improving GenEval image generation from 82% to 85%. Paper: arXiv 2606.27376.

Paper Overview

Field: Computer Vision Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawkar Published: 2026-06-27 arXiv: 2606.27376

Abstract

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask whether a unified LMM can improve both abilities autonomously using only unlabeled images. We propose a self-evolving training framework with three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that synthesizes images. Training uses only self-derived consistency signals, without human annotations, preference labels, or task-trained external reward/judge models.

To stabilize learning, we introduce Solver Token Entropy (STE), a continuous difficulty signal based on token-level prediction uncertainty. For image generation, we design a multi-scale internal evaluation scheme that combines question-answering fidelity scoring with cycle-consistency captioning.

The framework is applied to BLIP3o, BAGEL, and VARGPT-v1.1 architectures while keeping the same role decomposition and reward logic. Across eight understanding benchmarks, our method consistently outperforms the base models. On BAGEL, we achieve a +3.5% absolute gain on MMMU and improve GenEval image generation performance from 82% to 85%.

---

*Auto-collected on 2026-06-27*

Tags

#multimodal-models#self-evolving-training#image-generation#visual-understanding#reinforcement-learning#llm#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208198