Summary
This paper proposes a self-evolving training framework that enables unified large multimodal models (LMMs) to improve both visual understanding and image generation autonomously, without curated post-training supervision such as human annotations, preference labels, or external reward models. The framework assigns three internal roles to a single unified LMM: a Proposer that generates visual questions from unlabeled images, a Solver that answers and evaluates those questions, and a Generator that creates images. Improvement is driven by self-consistency rewards, allowing the model to supervise its own learning using only unlabeled images. The work, posted on arXiv (2606.27376), targets the computer vision research area and addresses the dependency of unified LMMs on costly curated supervision. The abstract was shared on zhichai.net, a Chinese tech forum, as part of an automated paper collection dated 2026-06-27.
Paper Overview
Research area: Computer Vision (CV)
Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar
Published: 2026-06-27
arXiv: 2606.27376
Abstract
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. This paper asks whether a unified LMM can improve both abilities autonomously using only unlabeled images.
The authors propose a self-evolving training framework built on three internal roles within the model:
- Proposer — generates visual questions from unlabeled images.
- Solver — answers and evaluates those questions.
- Generator — creates images.
The framework relies on self-consistency rewards to drive improvement, enabling the model to supervise its own training without external annotation or reward models.
Notes
- Forum post tags: paper, arXiv, CV
- Automatically collected on 2026-06-27
- Source excerpt is truncated; see the arXiv link above for the full paper.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178208189