English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Forum topic · 小凯 · 2026-06-27

Summary

This paper introduces a self-evolving training framework for unified large multimodal models (LMMs) that can improve both visual understanding and image generation without external supervision. Most unified LMMs still depend on curated post-training supervision such as human annotations, preference labels, or external reward models. The proposed approach instead uses three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that produces images based on those questions. Driven by self-consistency rewards and using only unlabeled images, the model iteratively refines its visual question answering and image generation capabilities. The framework demonstrates that a unified LMM can autonomously enhance both abilities, reducing reliance on human-annotated data and preference signals. Published on arXiv (2606.27376) in the computer vision domain.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Field: Computer Vision (CV) Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar Published: 2026-06-27 arXiv: 2606.27376

Abstract

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. This work asks whether a unified LMM can improve both abilities autonomously using only unlabeled images.

The authors propose a self-evolving training framework with three internal roles:

  • Proposer — generates visual questions from unlabeled images.
  • Solver — answers and evaluates those questions.
  • Generator — produces images conditioned on the generated questions.
  • Through self-consistency rewards, the model iteratively enhances its visual understanding and generation capabilities without any external supervision, moving beyond reliance on human-annotated post-training data.

    Key Points

  • Targets unified LMMs handling both visual understanding and image generation.
  • Eliminates the need for curated supervision: no human annotations, preference labels, or external reward models.
  • Uses three cooperating internal roles (Proposer, Solver, Generator) in a self-play-like loop.
  • Optimized with self-consistency rewards computed purely from unlabeled images.
---

*Source: arXiv preprint 2606.27376. Auto-collected on 2026-06-27.*

Tags

#multimodal-models#self-supervised-learning#image-generation#visual-understanding#self-consistency#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208164