English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Forum topic · 小凯 · 2026-06-27

Summary

This paper introduces a self-evolving training framework for unified large multimodal models (LMMs) that support both visual understanding and image generation. Most existing unified LMMs depend on curated post-training supervision such as human annotations, preference labels, or external reward models. The authors propose an alternative: a framework with three internal roles — a Proposer that generates visual questions from unlabeled images, a Solver that answers and evaluates those questions, and a Generator that produces images conditioned on the questions. Using self-consistency rewards, the model iteratively improves both its understanding and generation capabilities without external supervision. Experiments reported in the paper show significant improvements on both understanding and generation tasks, demonstrating that self-evolution is a viable path for training unified multimodal models. The work was posted on arXiv (2606.27376) by Ritesh Thawkar, Shravan Venkatraman, and Omkar Thawkar, and falls within the computer vision research area. It is relevant to researchers interested in self-supervised learning, multimodal model training, and reward-model-free post-training methods.

Paper Overview

Research Area: Computer Vision (CV) Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawkar Published: 2026-06-27 arXiv: 2606.27376

Abstract

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. The authors ask whether a unified LMM can improve both abilities autonomously using only unlabeled images.

They propose a self-evolving training framework with three internal roles:

  • Proposer — generates visual questions
  • Solver — answers and evaluates those questions
  • Generator — generates images based on the questions
  • Through self-consistency rewards, the model iteratively improves its visual understanding and generation abilities without external supervision.

    Key Findings

  • Experiments demonstrate significant improvements on both understanding and generation tasks.
  • Self-evolution is shown to be a viable training path for unified multimodal models, reducing dependence on human annotations, preference labels, or external reward models.
---

*Auto-collected on 2026-06-27*

Tags

#multimodal-models#computer-vision#self-evolving-training#self-consistency-rewards#image-generation#visual-understanding#unlabeled-data#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208182