English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AlphaGRPO: Group Relative Policy Optimization for Self-Reflective Multimodal Generation in UMMs

Forum topic · 小凯 · 2026-05-14

Summary

AlphaGRPO is a novel framework applying Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models (UMMs), enhancing multimodal generation without an additional cold-start stage. The approach enables two advanced reasoning capabilities: Reasoning Text-to-Image Generation, where the model infers implicit user intent, and Self-Reflective Refinement, where it autonomously diagnoses and corrects misalignment in generated outputs. To provide stable supervision for real-world multimodal generation, the authors introduce Decompositional Verifiable Reward (DVReward), which uses an LLM to decompose complex user requests into atomic, verifiable semantic and quality questions evaluated by a general MLLM, yielding reliable and interpretable feedback instead of a holistic scalar reward. Experiments show robust improvements on multimodal generation benchmarks including GenEval, TIIF-Bench, DPG-Bench, and WISE, plus significant gains on GEdit editing tasks without any editing-specific training. Paper: arXiv 2605.12495 by Runhui Huang, Jie Wu, Rui Yang, Zhe Liu, and Hengshuang Zhao.

论文概要

Research Area: Computer Vision Authors: Runhui Huang, Jie Wu, Rui Yang, Zhe Liu, Hengshuang Zhao Published: 2026-05-12 arXiv: 2605.12495

Overview

This paper proposes AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models (UMMs) to enhance multimodal generation capabilities without an additional cold-start stage.

The approach unlocks the model's intrinsic potential for advanced reasoning tasks:

  • Reasoning Text-to-Image Generation — the model actively infers implicit user intents rather than following prompts literally.
  • Self-Reflective Refinement — the model autonomously diagnoses and corrects misalignments in its generated outputs.
  • DVReward

    To address the challenge of providing stable supervision for real-world multimodal generation, the authors introduce the Decompositional Verifiable Reward (DVReward). Unlike holistic scalar rewards, DVReward utilizes an LLM to decompose complex user requests into atomic, verifiable semantic and quality questions. These are then evaluated by a general-purpose MLLM to provide reliable and interpretable feedback signals for reinforcement learning.

    Results

    Extensive experiments demonstrate that AlphaGRPO achieves robust improvements on multimodal generation benchmarks including GenEval, TIIF-Bench, DPG-Bench, and WISE. It also delivers significant gains on GEdit editing tasks — notably without any training on editing tasks.

    Links

  • arXiv: https://arxiv.org/abs/2605.12495
---

*Auto-collected on 2026-05-14*

Tags

#alphagrpO#grpo#unified-multimodal-models#text-to-image-generation#reinforcement-learning#dvreward#ar-diffusion#self-reflective-refinement

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620001