English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability of Samplers

Forum topic · 小凯 · 2026-07-12

Summary

A new arXiv paper (2607.08757) by Yiwei Zhou shows that small forward-marginal score error does not guarantee numerical stability of diffusion model samplers. The author constructs a single smooth score field with arbitrarily small forward-marginal L2 error whose learned reverse-time process is nonexplosive and arbitrarily close to the exact reverse-time process in path-space total variation, yet its Euler–Maruyama discretizations converge in probability while every positive moment diverges—weak convergence can hold even as every Wasserstein distance Wp (p≥1) diverges. The failure can occur even within one fixed finite neural architecture, via a family of bounded, globally Lipschitz denoisers. A positive result for compactly supported data shows that projecting the learned denoiser onto a known bounded closed convex set preserves pointwise accuracy, yields grid-uniform moment bounds, and produces Wasserstein convergence under mild local regularity. Experiments with small fixed DiT-style networks confirm the phenomenon and the effectiveness of projection.

Paper Overview

Field: Machine Learning Author: Yiwei Zhou Published: 2026-07-09 arXiv: 2607.08757

Summary

Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. This paper proves that small forward-marginal error does not guarantee numerical stability.

Key Findings

  • The author constructs a single smooth score field with arbitrarily small forward-marginal \(L^2\) error.
  • The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation.
  • Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance \(W_p\), \(p \ge 1\), diverges.
  • The same failure can occur within one fixed finite neural architecture: there exists a family of bounded, globally Lipschitz denoisers whose forward-marginal error and path-space total variation distance both tend to zero, while their Euler--Maruyama endpoints diverge in every \(W_p\).
  • Positive Result for Compactly Supported Data

    For compactly supported data, a simple fix is given: projecting the learned denoiser onto a known bounded closed convex set containing the support

  • preserves pointwise accuracy,
  • gives grid-uniform moment bounds, and
  • yields Wasserstein convergence under mild local regularity.

Experiments

Experiments using small fixed DiT-style networks show large growth along rare numerical trajectories; denoiser projection suppresses this growth while overall trajectory error remains small.

--- *Auto-collected on 2026-07-12*

Tags

#diffusion-models#score-matching#numerical-stability#euler-maruyama#wasserstein-distance#sampling#theory#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379393