Paper Overview
- Research area: Computer Vision (CV)
- Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar
- Published: 2026-06-27
- arXiv: 2606.27376
- Proposer — generates visual questions
- Solver — answers and evaluates those questions
- Generator — produces images based on the questions
Abstract
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. This paper asks whether a unified LMM can improve both abilities autonomously using only unlabeled images.The authors propose a self-evolving training framework with three internal roles:
Key Takeaways
1. Removes reliance on human annotations, preference labels, and external reward models. 2. Uses only unlabeled images as training signal. 3. Couples understanding and generation improvement within a single unified model via role-based self-play.--- *Auto-collected on 2026-06-27*