[论文] Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

论文概要 研究领域: cs.CV 作者: Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao 发布时间: 20…

论文概要

研究领域: cs.CV 作者: Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao 发布时间: 2026-09-08 arXiv: 2609.09123

中文摘要

自回归(AR)视频扩散模型在实时视频生成方面展现出巨大潜力。近期方法通过分布匹配蒸馏(DMD)将预训练的双向视频扩散模型蒸馏为因果AR学生模型,但生成的视频常遭受过饱和和过平滑问题,导致视觉质量和真实感有限。关键促成因素是DMD中反向KL目标的模式寻找行为,这可能导致学生分布坍缩到教师分布的少数模式上。为解决此问题,我们提出Mask Forcing,一种双噪声掩码展开策略,扰动AR学生自展开以缓解反向KL模式寻找引起的模式坍缩。核心思想是在AR扩散蒸馏的自展开过程中,沿空间和时间轴通过随机掩码将更干净的信号注入噪声展开输入。这种扰动鼓励学生展开探索教师分布的更多区域,使DMD能够提供超越学生已覆盖模式的学习信号。此外,更干净的token作为去噪指导作用于其他更噪声的token,改进了中间展开预测并减少了误差累积。大量实验表明,我们的方法以更高视觉质量有效改进了多种AR视频扩散蒸馏方法,且无需引入真实视频数据或额外的后训练阶段。

原文摘要

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.


*自动采集于 2026-09-10*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens