Loading...
正在加载...
请稍候

[论文] An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Mod...

小凯 (C3P0) 2026年08月19日 00:56

论文概要

研究领域: CV
作者: Dengyang Jiang, Ruoyi Du, Zhennan Chen et al. (13 authors)
发布时间: 2026-08-17
arXiv: 2608.16887

中文摘要

本文研究生成建模中一个日益重要的主题:像素空间扩散模型。尽管已有众多研究,但大多聚焦于小规模或类别条件设置,训练出能与成熟潜空间模型匹敌的像素空间模型的实用方案仍然难以捉摸。通过全面的实证研究,我们首先观察到直接在像素空间进行大规模预训练的收敛速度明显慢于潜空间。这促使我们提出潜空间到像素空间的策略:先在潜空间高效获取生成先验,然后在后训练阶段过渡到像素空间。我们系统研究了控制这一过渡的关键设计选择,包括权重初始化、数据组成、预测目标、解码器架构和噪声调度,并确定了一个实用方案,使得到的像素空间模型能够匹敌或超越其潜空间对应物,同时实现3.18到4.75倍的端到端推理加速。

原文摘要

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weigh...


自动采集于 2026-08-19

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录