[论文] Adversarial Training for Pixel Diffusion
研究领域: CV 作者: Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen 发布时间: 2026-09-29 arXiv: 2609.38170
论文概要
研究领域: CV 作者: Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen 发布时间: 2026-09-29 arXiv: 2609.38170
中文摘要
像素空间扩散模型直接生成RGB图像,避免了自编码器的瓶颈,但其输出仍未充分再现细粒度自然图像统计特征。本文表明对抗学习为这一缺陷提供了一种有效的训练后修正方法。从预训练模型出发,我们保留其原始的扩散或流匹配目标,并在非高噪声时间步上对预测输出添加对抗损失,模型架构和采样过程均保持不变。据我们所知,这是首个对像素扩散模型进行系统性对抗训练后研究的工作。在两种像素骨干网络上,该方法同时改善了分布保真度、覆盖率、提示对齐度和感知质量。我们进一步探究了其内在机理:频带和幂律分析表明,原始模型系统性地产出不足的自然图像高频内容,而对抗训练后恢复了这一缺失的频谱功率。相比之下,感知损失也能增加高频内容,但会牺牲分布保真度和提示对齐度。最近邻、召回率和匹配的无GAN SFT对照实验进一步排除了记忆化、模式丢弃和额外优化作为简单解释。最后,我们考察了这一效应的边界:在测试的潜在扩散配置下,相同程序未产生可比的联合改进,且几乎未增加解码后的高频功率。这些结果表明,直接输出访问被修正的图像统计特征是决定对抗训练后成功与否的关键因素。
原文摘要
Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. ...
*自动采集于 2026-10-01*
#论文 #arXiv #CV #小凯