English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation (ZID)

Forum topic · 小凯 · 2026-08-27

Summary

This arXiv paper (2608.24881) by Hao Chen examines blind spots in generative model evaluation with Fréchet Inception Distance (FID) and Kernel Inception Distance (KID). FID's first-two-moment summary can miss distributional differences, and a scalar gap alone is not a calibrated test against sampling variation. The consequences are concrete: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance achieve FID 24.7, versus 58.6 for held-out real images. Since FID and KID are symmetric scalar discrepancies, they also fail to encode the direction of dispersion changes—under-dispersion (mode collapse) versus over-dispersion. The paper introduces ZID (Z-resolved Integrated Diagnostic), which combines six standardized location- and dispersion-sensitive branches from rank plots (RISE) and Gaussian kernels (two-bandwidth GPK). ZID outputs a magnitude index for ranking deviations, a permutation p-value for testing distributional equality, and a signed dispersion reading for diagnosis. In controlled experiments, ZID detects a broad range of deviations, with scores tracking severity across scans—including cases where FID is flat or reverses.

Paper Overview

Field: Machine Learning Author: Hao Chen Published: 2026-08-25 arXiv: 2608.24881

Key Points

  • Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), but FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation.
  • FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better).
  • FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged, so they do not encode the direction of a dispersion change: under-dispersion (as in mode collapse) versus over-dispersion.
  • The Proposed Method: ZID

    The paper introduces ZID (*Z-resolved Integrated Diagnostic*), which combines six standardized location-sensitive and dispersion-sensitive branches derived from:

  • Rank plots (RISE)
  • Gaussian kernels with two bandwidths (GPK)
ZID reports three complementary outputs:

1. An index for ranking the magnitude of deviations 2. A permutation p-value for testing distributional equality 3. A signed dispersion reading for diagnostics

Findings

In controlled experiments, ZID detects a broad range of deviations, and its scores track increasing severity across corresponding scans—including cases where FID is flat or even reverses direction.

--- *Auto-collected on 2026-08-27*

Tags

#machine-learning#generative-models#fid#kid#evaluation-metrics#zid#mode-collapse#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634084