Paper Overview
Field: Computer Vision (CV) Authors: Rogerio Guimaraes, Pietro Perona Published: 2025-07-27 arXiv: 2507.21744
Abstract (Translation)
Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under a black-box reward, but typically maintain a constant memory footprint throughout inference. The authors show that relaxing this constraint enables an underexplored inference-time scaling axis: by front-loading exploration, evaluating many seeds early, and pruning aggressively, a fixed compute budget can be used more effectively.
Method: Progressive Seed Pruning (PSP)
- Score intermediate denoised estimates of many candidate seeds early in the sampling process.
- Progressively narrow the candidate set, so only promising trajectories are fully denoised.
- Keep the total number of model evaluations fixed, reallocating compute toward the most promising trajectories.
- Evaluated on both diffusion and flow-matching backbones.
- PSP consistently improves reward-guided selection compared to baselines.
- At equal compute, PSP achieves higher GenEval scores (automated evaluation) and better human-rated prompt alignment than best-of-N, importance sampling, and tree-search baselines.
- arXiv: <https://arxiv.org/abs/2507.21744>