English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Data Geometry of Masking Diffusion: Certified-Optimal Schedules via Unmasking Growth Complexity

Forum topic · 小凯 · 2026-08-15

Summary

This arXiv paper (2608.13520) by Martin J. Wainwright studies masking diffusion models for discrete sampling and introduces a path-resolved measure of data geometry called unmasking growth complexity (UGC). Local increments of UGC directly control Kullback-Leibler (KL) discretization error, enabling a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes. In log-reveal-odds coordinates, this structure yields optimized single-block and multi-block schedules and quantifies the gains from adapting computational effort to data geometry. The paper shows UGC increments can be estimated from samples via KL increments along coupled reveal trajectories, leading to certified-optimal samplers that achieve a prescribed KL error with high probability at iteration complexity within a constant factor of the oracle procedure. The aggregate UGC mass connects to classical multivariate dependence measures, and in the fine-partition limit the squared integral of the square-root UGC density determines the sharp leading-order optimal Euler discretization error. Examples show substantial dimension-dependent gains, including Omega-tilde(sqrt(d)) improvements with a constant number of adaptively placed blocks.

Paper Overview

  • Field: Machine Learning
  • Author: Martin J. Wainwright
  • Published: 2026-08-13
  • arXiv: 2608.13520
  • Abstract

    We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the unmasking growth complexity (UGC). Its local increments directly control Kullback--Leibler (KL) discretization error, yielding a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes.

    In log-reveal-odds coordinates, this structure yields optimized single-block and multi-block schedules, and quantifies the gains from adapting computational effort to data geometry. Crucially, we show how UGC increments can be estimated from samples via KL increments along coupled reveal trajectories. This leads to certified-optimal samplers that achieve a prescribed KL error with high probability and iteration complexity within a constant factor of the corresponding oracle procedure.

    Collapsing the UGC path yields the aggregate UGC mass, which connects to classical multivariate dependence measures and complexity measures from previous analyses of discrete diffusion. In the fine-partition limit, the squared integral of the square-root UGC density determines the sharp leading-order optimal Euler discretization error. Examples exhibit substantial dimension-dependent gains over coarse schedules, including Ω̃(√d) improvements achievable with a constant number of adaptively placed blocks.

    Key Contributions

  • A path-resolved data geometry measure (UGC) whose local increments control KL discretization error
  • Unified analysis covering both Bernoulli-subset and fixed-cardinality unmasking schemes
  • Optimized single-block and multi-block schedules in log-reveal-odds coordinates
  • Sample-based estimation of UGC increments, enabling certified-optimal samplers with provable KL error guarantees
  • Connections between aggregate UGC mass and classical multivariate dependence measures
  • Sharp characterization of optimal Euler discretization error via the square-root UGC density in the fine-partition limit
  • Demonstrated Ω̃(√d) dimension-dependent gains using a constant number of adaptively placed blocks
---

*Auto-collected on 2026-08-15*

Tags

#machine-learning#diffusion-models#discrete-sampling#sampling-schedules#theory#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633508