Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Field: Computer Vision (CV) Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar Published: 2026-06-27 arXiv: 2606.27376
Abstract
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. This work asks whether a unified LMM can improve both abilities autonomously using only unlabeled images.
The authors propose a self-evolving training framework with three internal roles:
- Proposer — generates visual questions from unlabeled images.
- Solver — answers and evaluates those questions.
- Generator — produces images conditioned on the generated questions.
- Targets unified LMMs handling both visual understanding and image generation.
- Eliminates the need for curated supervision: no human annotations, preference labels, or external reward models.
- Uses three cooperating internal roles (Proposer, Solver, Generator) in a self-play-like loop.
- Optimized with self-consistency rewards computed purely from unlabeled images.
Through self-consistency rewards, the model iteratively enhances its visual understanding and generation capabilities without any external supervision, moving beyond reliance on human-annotated post-training data.
Key Points
*Source: arXiv preprint 2606.27376. Auto-collected on 2026-06-27.*