Paper Overview
- Field: Computer Vision (CV)
- Authors: Tianhong Li, Huiwen Chang, Kai Zhang, et al.
- arXiv: 2604.28190
Key Idea
The paper demonstrates that Fréchet Distance (FD) — long considered impractical as a training objective — can in fact be effectively optimized in representation spaces. The core insight is simple: decouple the population size for FD estimation (e.g., 50k samples) from the batch size for gradient computation (e.g., 1024). The authors call this approach FD-loss.
Findings
Optimizing FD-loss reveals several surprising results:
1. Post-training improves visual quality: Applying FD-loss to post-train a base generator in different representation spaces consistently improves visual quality. Using Inception feature space, a one-step generator achieves 0.72 FID on ImageNet 256x256.
2. Multi-step to one-step conversion: The same FD-loss can transform a multi-step generator into a powerful one-step generator — without teacher distillation, adversarial training, or per-sample objectives.
3. FID may misrank visual quality: Modern representations can produce visually better samples even when Inception FID is worse. This motivates FDr^k, a multi-representation evaluation metric.
Conclusion
The authors hope this work encourages further exploration of distribution distances in diverse representation spaces, both as training objectives and as evaluation metrics for generative models.
---
*Originally shared on zhichai.net, 2026-05-02.*