论文概要
研究领域: CV
作者: Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
发布时间: 2026-09-29
arXiv: 2609.38154
中文摘要
视频扩散模型越来越多地被开发为面向多样化下游任务的专用模型,这一开发通常包含一个蒸馏阶段——例如加速采样或改进长视频生成。这个阶段通常为每个专用模型重复进行。我们引入LongLive-Plug——一个一次蒸馏(once-for-all)框架,在基础模型上以LoRA形式学习可复用能力,实现对兼容下游模型的免训练、即插即用部署。这些能力包括单步分类器无关引导(CFG)、少步采样和自回归生成的长上下文纠错。即使下游模型添加条件分支或扩展输出通道,这些适配器仍然可复用。尽管训练时使用固定的引导尺度,我们的专用CFG LoRA通过推理权重提供文本引导控制。将其与少步LoRA结合,同时在下游任务上保持少步生成和CFG可控性。我们在三个骨干家族、八个任务类别(包括世界建模、机器人、编辑和多模态生成)的54个下游模型上验证了免训练部署。每种能力可以为每个骨干家族蒸馏一次,无需按目标重新训练即可复用。
原文摘要
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, o...
自动采集于 2026-10-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。