English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Representation Fréchet Loss for Visual Generation: Turning Fréchet Distance into a Training Objective

Forum topic · 小凯 · 2026-05-02

Summary

This paper (arXiv:2604.28190, Tianhong Li, Huiwen Chang, Kai Zhang, et al.) shows that Fréchet Distance (FD), long considered impractical as a training objective, can be effectively optimized in representation spaces. The key idea, called FD-loss, is to decouple the population size for FD estimation (e.g., 50k samples) from the batch size used for gradient computation (e.g., 1024). The authors report three surprising findings. First, post-training a base generator with FD-loss in different representation spaces consistently improves visual quality; a one-step generator reaches 0.72 FID on ImageNet 256x256 in Inception feature space. Second, the same FD-loss can convert multi-step generators into powerful one-step generators without teacher distillation, adversarial training, or per-sample objectives. Third, FID may rank visual quality incorrectly: modern representations can yield better samples despite worse Inception FID, motivating FDr^k, a multi-representation evaluation metric. The work encourages exploring distribution distances across diverse representation spaces as both training objectives and evaluation metrics for generative models.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Tianhong Li, Huiwen Chang, Kai Zhang, et al.
  • arXiv: 2604.28190

Key Idea

The paper demonstrates that Fréchet Distance (FD) — long considered impractical as a training objective — can in fact be effectively optimized in representation spaces. The core insight is simple: decouple the population size for FD estimation (e.g., 50k samples) from the batch size for gradient computation (e.g., 1024). The authors call this approach FD-loss.

Findings

Optimizing FD-loss reveals several surprising results:

1. Post-training improves visual quality: Applying FD-loss to post-train a base generator in different representation spaces consistently improves visual quality. Using Inception feature space, a one-step generator achieves 0.72 FID on ImageNet 256x256.

2. Multi-step to one-step conversion: The same FD-loss can transform a multi-step generator into a powerful one-step generator — without teacher distillation, adversarial training, or per-sample objectives.

3. FID may misrank visual quality: Modern representations can produce visually better samples even when Inception FID is worse. This motivates FDr^k, a multi-representation evaluation metric.

Conclusion

The authors hope this work encourages further exploration of distribution distances in diverse representation spaces, both as training objectives and as evaluation metrics for generative models.

---

*Originally shared on zhichai.net, 2026-05-02.*

Tags

#computer-vision#generative-models#frechet-distance#fid#one-step-generation#arxiv-paper#training-objective#image-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619029