论文概要
研究领域: CV
作者: Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Shaoteng Liu, Lehan Yang, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen
发布时间: 2026-09-29
arXiv: 2609.38156
中文摘要
分布匹配蒸馏(DMD)为少步扩散生成提供了一个通用框架,但其现代文本到图像实例主要围绕潜在扩散开发,因此忽略了原生RGB空间的关键特性和设计机会。我们重新审视了像素空间教师的两个DMD接口。在教师匹配侧,诊断显示低噪声RGB匹配由局部纹理线索主导,这促使我们采用固定的高噪声匹配频带。在真实数据侧,原生干净RGB输出允许来自外部视觉表示的引导,而无需遍历解码器或共享沉重的假分数评判器。DINO-Adv将此评判器从对抗梯度路径中移除,提供局部参数化补丁引导。对于分布级引导,我们提出AF-Loss——一种无参数的辅助语义分布场目标,专为文本到图像DMD设计。它在共享的DINOv2空间中作用于分离的滚动真实和生成支撑集,同时保留提示条件的教师监督。AF-Loss不增加可学习参数或推理时计算。这些设计共同构成了DMA²。在DPG-Bench、GenEval、VQAScore和COCO30K上,四步DMA²学生模型表现优于25步教师模型和已评估的少步蒸馏方法。
原文摘要
Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidan...
自动采集于 2026-10-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。