[论文] $λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Bu...

研究领域: ML 作者: Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa 发布时间: 2026-09-18 arXiv: 2609.22041

论文概要

研究领域: ML 作者: Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa 发布时间: 2026-09-18 arXiv: 2609.22041

中文摘要

强化学习正越来越多地用于让图像生成器与奖励信号对齐。Flow-GRPO 最近将这一范式扩展到 flow-matching 模型,将去噪采样器视为可从奖励反馈中优化的随机策略。该设定下的训练以一种多步去噪特有的方式不稳定:策略更新在各去噪步骤间系统性变化,重要性比率漂移到 1 以下、越发分散、以不同比率被裁剪,并在训练后期留下更少的可用样本。先前工作将这些效应视为独立的失效模式,并各用一个人工调谐的稳定器应对。而我们证明,它们都源于一个单一的逐步量,称之为路径方差(path variance)。该量由采样器的高斯转移核精确决定,且可在训练中廉价估计。这把不稳定性重新框架化为一种可测量、可预算的资源,而非一堆待修复的症状。我们的方法 λ-Controlled GRPO 从这个预测律而非嘈杂的经验统计中校准重要性比率行为,并按各去噪步骤的预测成本分配梯度努力。控制更新的两个尺度由标准策略选择确定,而非作为自由调参引入。在一个文本到图像模型、两种奖励设置下——由 OCR 评分的难渲染目标文本、由偏好模型评分的人类偏好匹配——λ-Controlled GRPO 在文本准确度和偏好奖励上均优于最强的经验稳定器。它还将后期步骤的路径方差保持在预算内,而这正是基线系统性超支之处。最终得到一个由自身转移律校准的 Flow-GRPO 更新,而非在不稳定出现后再进行稳定化。

原文摘要

Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a hand-tuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly ...


*自动采集于 2026-09-22*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens