English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SLAS: Super-Linear Advantage Shaping for Reinforcement Post-Training of Text-to-Image Models

Forum topic · 小凯 · 2026-05-13

Summary

This arXiv paper (2505.07245) proposes Super-Linear Advantage Shaping (SLAS), a new method for reinforcement learning post-training of text-to-image (T2I) models. Building on Group Relative Policy Optimization (GRPO), the authors identify that normalization can cause miscalibration and that reward hacking—where models exploit biases in imperfect reward functions—remains a key problem. SLAS revisits the functional update from an information-geometric perspective, extending the Fisher-Rao information metric with advantage-dependent weighting. This reshapes the local policy space with a nonlinear geometry: constraints are relaxed along high-advantage directions to amplify informative updates, and tightened in low-advantage regions to suppress spurious gradients. Batch-level normalization stabilizes training across reward scales. Evaluations show SLAS consistently outperforms the DanceGRPO baseline across multiple backbones and benchmarks, delivering faster training dynamics, improved out-of-domain performance on GenEval and UniGenBench++, and stronger model-scaling robustness while mitigating reward hacking and preserving semantic and compositional fidelity in generation.

Overview

  • Field: Computer Vision (CV)
  • Authors: Haoyuan Sun, Jing Wang, Yuxin Song
  • Published: 2025-05-09
  • arXiv: 2505.07245
  • Abstract (translated)

    Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for further advancement of text-to-image (T2I) models. However, these methods are often prone to reward hacking, wherein models exploit biases in imperfect reward functions rather than yielding genuine performance gains. In this work, the authors identify that normalization could lead to miscalibration, and that directly removing the prompt-level standard deviation term yields an optimal policy ascent direction that is linear in the advantage but still limits the separation of genuine signals from noise.

    To address these issues, they propose Super-Linear Advantage Shaping (SLAS) by revisiting the functional update from an information-geometric perspective. By extending the Fisher-Rao information metric with advantage-dependent weighting, SLAS introduces a nonlinear geometric structure that reshapes the local policy space. The design relaxes constraints along high-advantage directions to amplify informative updates, while tightening constraints in low-advantage regions to suppress spurious gradients. In addition, batch-level normalization is applied to stabilize training across different reward scales.

    Key Results

  • Extensive evaluations show SLAS consistently surpasses the DanceGRPO baseline across multiple backbones and benchmarks.
  • It produces faster training dynamics and improved out-of-domain performance on GenEval and UniGenBench++.
  • SLAS offers enhanced robustness to model scaling while mitigating reward hacking and preserving semantic and compositional fidelity in generation.

Tags

#text-to-image#reinforcement-learning#grpo#reward-hacking#information-geometry#diffusion-models#post-training#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619912