English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Forum topic · 小凯 · 2026-06-27

Summary

This paper introduces a self-evolving training framework that enables unified large multimodal models (LMMs) to improve both visual understanding and image generation autonomously, without curated post-training supervision such as human annotations, preference labels, or external reward models. The approach assigns three internal roles within a single model: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that produces images based on those questions. Training is driven by self-consistency rewards computed only from unlabeled images, allowing the model to iteratively refine its understanding and generation capabilities in a fully self-supervised loop. The work addresses the dependence of most unified LMMs on curated supervision and demonstrates a path toward continuous self-improvement using unlabeled image data alone. Authored by Ritesh Thawkar, Shravan Venkatraman, and Omkar Thawakar, the paper (arXiv:2606.27376) falls in the computer vision research area.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar
  • Published: 2026-06-27
  • arXiv: 2606.27376
  • Abstract

    Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. This paper asks whether a unified LMM can improve both abilities autonomously using only unlabeled images.

    The authors propose a self-evolving training framework with three internal roles:

  • Proposer — generates visual questions
  • Solver — answers and evaluates those questions
  • Generator — produces images based on the questions
Through self-consistency rewards, the model iteratively improves its visual understanding and generation capabilities without external supervision.

Key Takeaways

1. Removes reliance on human annotations, preference labels, and external reward models. 2. Uses only unlabeled images as training signal. 3. Couples understanding and generation improvement within a single unified model via role-based self-play.

--- *Auto-collected on 2026-06-27*

Tags

#multimodal-models#self-supervised-learning#image-generation#computer-vision#self-consistency#lmm#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208168