English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models: A Latent-to-Pixel Recipe

Forum topic · 小凯 · 2026-08-19

Summary

This arXiv paper (2608.16887) by Dengyang Jiang, Ruoyi Du, Zhennan Chen et al. investigates pixel-space diffusion models for text-to-image generation. The authors observe that direct large-scale pre-training in pixel space converges substantially slower than in latent space, which has limited practical pixel-space models to small-scale or class-conditional settings. They propose a latent-to-pixel strategy: efficiently acquire generative priors in latent space, then transition to pixel space during post-training. The study systematically examines key design choices for this transition, including weight initialization, data composition, prediction objective, decoder architecture, and noise schedules. The resulting recipe enables pixel-space models to match or exceed latent-space counterparts while achieving 3.18x to 4.75x end-to-end inference speedup.

Overview

Field: Computer Vision (CV) Authors: Dengyang Jiang, Ruoyi Du, Zhennan Chen et al. (13 authors) Published: 2026-08-17 arXiv: 2608.16887

Key Points

  • The paper studies pixel-space diffusion models for text-to-image generation. While many studies explore this topic, most focus on small-scale or class-conditional settings, leaving open a practical recipe for training pixel-space models that rival well-established latent-space counterparts.
  • Through a comprehensive empirical study, the authors observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space.
  • This motivates a latent-to-pixel strategy: efficiently acquire generative priors in latent space first, then transition to pixel space during post-training.
  • The transition's key design choices are systematically investigated, including:
  • Weight initialization
  • Data composition
  • Prediction objective
  • Decoder architecture
  • Noise scheduling
  • The identified practical recipe yields pixel-space models that match or exceed their latent-space counterparts while delivering 3.18x to 4.75x end-to-end inference acceleration.

Original Abstract (excerpt)

> This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space...

*Auto-collected on 2026-08-19*

Tags

#diffusion-models#text-to-image#pixel-space#latent-space#generative-modeling#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633634