Summary
This paper introduces a self-evolving training framework for unified large multimodal models (LMMs) that support both visual understanding and image generation. Most existing unified LMMs depend on curated post-training supervision such as human annotations, preference labels, or external reward models. The authors propose an alternative: a framework with three internal roles — a Proposer that generates visual questions from unlabeled images, a Solver that answers and evaluates those questions, and a Generator that produces images conditioned on the questions. Using self-consistency rewards, the model iteratively improves both its understanding and generation capabilities without external supervision. Experiments reported in the paper show significant improvements on both understanding and generation tasks, demonstrating that self-evolution is a viable path for training unified multimodal models. The work was posted on arXiv (2606.27376) by Ritesh Thawkar, Shravan Venkatraman, and Omkar Thawkar, and falls within the computer vision research area. It is relevant to researchers interested in self-supervised learning, multimodal model training, and reward-model-free post-training methods.
Paper Overview
Research Area: Computer Vision (CV)
Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawkar
Published: 2026-06-27
arXiv: 2606.27376
Abstract
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. The authors ask whether a unified LMM can improve both abilities autonomously using only unlabeled images.
They propose a self-evolving training framework with three internal roles:
- Proposer — generates visual questions
- Solver — answers and evaluates those questions
- Generator — generates images based on the questions
Through self-consistency rewards, the model iteratively improves its visual understanding and generation abilities without external supervision.
Key Findings
- Experiments demonstrate significant improvements on both understanding and generation tasks.
- Self-evolution is shown to be a viable training path for unified multimodal models, reducing dependence on human annotations, preference labels, or external reward models.
---
*Auto-collected on 2026-06-27*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178208182