English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Forum topic · 小凯 · 2026-07-18

Summary

MeanFlowNFT is a reinforcement learning framework that extends forward-process RL tuning to MeanFlow generators, which achieve fast few-step sampling by predicting average velocities over time intervals. Existing methods like DiffusionNFT optimize instantaneous velocities, while MeanFlow samples with average velocities, leaving a gap for applying such RL approaches. MeanFlowNFT bridges this by constructing an induced instantaneous-velocity predictor inspired by the MeanFlow identity, applying the DiffusionNFT objective to it, and retaining average-velocity sampling for few-step generation. The method inherits DiffusionNFT's strict policy improvement guarantees. Experiments on image and video generation show consistent improvements over baselines, outperforming prior few-step RL-tuned generators on 6 of 8 metrics for SD3.5-M, and even surpassing multi-step RL-tuned diffusion models: 4-step MeanFlowNFT reaches a VBench score of 84.33 on Wan 2.1, exceeding 50-step LongCat-Video RL (82.57). Paper: arXiv 2607.15273.

Paper Overview

Field: Computer Vision (CV) Authors: Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang Posted: 2026-07-16 arXiv: 2607.15273

Abstract (Translated)

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive candidates for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that requires neither reverse-process trajectories nor likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored: DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities.

To bridge this gap, the authors introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, they construct an induced instantaneous-velocity predictor. The DiffusionNFT objective is applied to this predictor, making reward optimization for MeanFlow well-defined. Sampling still relies on average velocities, preserving MeanFlow's fast few-step generation advantage. The paper further proves that MeanFlowNFT inherits DiffusionNFT's strict policy improvement guarantees.

Results

Experiments on image and video generation show MeanFlowNFT consistently improves over baselines:

  • Outperforms prior state-of-the-art few-step RL-tuned generators on 6 of 8 metrics for SD3.5-M.
  • With only a few sampling steps, it surpasses even multi-step RL-tuned diffusion models.
  • On Wan 2.1, 4-step MeanFlowNFT achieves a VBench score of 84.33, exceeding 50-step LongCat-Video RL (82.57).
  • Links

  • arXiv: <https://arxiv.org/abs/2607.15273>
--- *Auto-collected on 2026-07-18*

Tags

#meanflownft#reinforcement-learning#diffusion-models#few-step-generation#flow-matching#video-generation#image-generation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433579