[论文] LongLive-Plug: Once-for-All Distillation for Video Generation

研究领域: CV 作者: Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen 发布时…

论文概要

研究领域: CV 作者: Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen 发布时间: 2026-09-29 arXiv: 2609.38154

中文摘要

视频扩散模型越来越多地被开发为面向多样化下游任务的专用模型,这一开发通常包含一个蒸馏阶段——例如加速采样或改进长视频生成。这个阶段通常为每个专用模型重复进行。我们引入LongLive-Plug——一个一次蒸馏(once-for-all)框架,在基础模型上以LoRA形式学习可复用能力,实现对兼容下游模型的免训练、即插即用部署。这些能力包括单步分类器无关引导(CFG)、少步采样和自回归生成的长上下文纠错。即使下游模型添加条件分支或扩展输出通道,这些适配器仍然可复用。尽管训练时使用固定的引导尺度,我们的专用CFG LoRA通过推理权重提供文本引导控制。将其与少步LoRA结合,同时在下游任务上保持少步生成和CFG可控性。我们在三个骨干家族、八个任务类别(包括世界建模、机器人、编辑和多模态生成)的54个下游模型上验证了免训练部署。每种能力可以为每个骨干家族蒸馏一次,无需按目标重新训练即可复用。

原文摘要

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, o...


*自动采集于 2026-10-01*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens