Paper Overview
- Field: Computer Vision (CV)
- Authors: Rogerio Guimaraes, Pietro Perona
- Published: 2026-07-24
- arXiv: 2507.19318
- Seed choice strongly affects final generation quality, making seed search a natural scaling lever.
- PSP trades a larger memory footprint early in inference for better sample efficiency, keeping total model evaluations fixed.
- Progressive pruning avoids wasting full denoising passes on unpromising seeds.
- PSP outperforms best-of-N, importance sampling, and tree-search baselines on GenEval and human evaluations.
Abstract (Full Translation)
Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under a black-box reward, but typically maintain a constant memory footprint throughout inference. The authors show that relaxing this constraint enables an underexplored inference-time scaling axis: by front-loading exploration, evaluating many seeds early, and pruning aggressively, a fixed compute budget can be used more effectively.
Progressive Seed Pruning (PSP) scores intermediate denoised estimates and progressively narrows the candidate set so that only promising trajectories are fully denoised, while keeping the total number of model evaluations fixed. Across diffusion and flow-matching backbones, PSP consistently improves over reward-guided selection, achieving higher GenEval scores (automatic evaluation) and better human-rated prompt alignment than best-of-N, importance sampling, and tree-search baselines at matched compute.
Key Takeaways
*Auto-collected on 2026-07-25*