Summary
This paper introduces a training-free self-guidance mechanism to mitigate diversity collapse in state-of-the-art flow models used for text- and image-conditioned image generation. When generating multiple samples under the same conditioning, flow models tend to collapse toward a single mode, producing visually similar outputs. Existing remedies either apply latent guidance, which has limited effectiveness, or perform sample selection using external reward models, which adds significant inference-time overhead. The proposed feature self-guidance approach instead leverages the model's internal features to encourage diverse outputs without retraining or auxiliary reward networks, offering an efficient solution at inference time. Authored by Pradhaan S Bhat, Rishubh Parihar, and Abhijnya Bhat in the computer vision field, the work is available on arXiv (2606.27371).
Paper Overview
Research Area: Computer Vision (CV)
Authors: Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat
Published: 2026-06-27
arXiv: 2606.27371
Abstract
State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples under the same conditioning — repeated generations tend to settle at a single mode, producing similar images.
Existing methods address this issue via either:
- Latent guidance, which has limited effectiveness
- Sample selection, which relies on external reward models that incur significant inference-time overhead
In this work, the authors introduce an efficient, training-free self-guidance mechanism to mitigate diversity collapse. Instead of relying on external reward models or retraining, the method guides generation using the model's own internal features, promoting diverse outputs while avoiding added inference overhead.
Links
- arXiv: https://arxiv.org/abs/2606.27371
---
*Auto-collected on 2026-06-27.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178208193