[论文] Learning Native Reflection in Unified Models with Interleaved Reinforc...
研究领域: CV 作者: Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu 发布时间: 2026-09-28 arXiv: 2609.35767
论文概要
研究领域: CV 作者: Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu 发布时间: 2026-09-28 arXiv: 2609.35767
中文摘要
统一多模态模型既能看图也能生成图像,因此原则上它们可以修复自己的生成结果:诊断图像哪里错了,修改它,观察结果,再诊断。修改是否有用只有在渲染后才知道,因此反思文本和图像生成必须在整个循环中联合学习。在反思轨迹上进行有监督微调(SFT)可以给一个冷启动,但找不到高成功率的修复路径;而只优化渲染器或只优化一个头的朴素 RL 则让大部分收益未被发掘。我们提出 UMM-Reflection,将强化学习(RL)应用于统一模型内部完整反思轨迹的训练:兄弟轨迹共享同一初始图像,因此组相对优势可以比较不同的反思策略;轨迹级优势同时更新反思 token 和基于流的修改,避免了逐轮信用分配的组合爆炸。与单轮编辑或带外部评论器的流水线不同,信用可以跨轮流动并流向同一模型的两个角色,推理时无需任何验证器。在 BAGEL 上,UMM-Reflection 将 GenEval 比 SFT 提高了 12.05 分,且收益迁移到 WISE(+10.97)、OneIG-Bench(+3.48)和 T2I-CompBench++(+4.63),这些均未在训练中使用。
原文摘要
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares r...
*自动采集于 2026-09-30*
#论文 #arXiv #CV #小凯