论文概要
研究领域: CV
作者: Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam, Donghoon Lee, Hyunsung Go, Yeonkyeong Lee, Hansaem Kim, Seungryong Kim
发布时间: 2026-10-07
arXiv: 2610.10524
中文摘要
高压缩率的视频自编码器为加速视频扩散模型提供了一种有效途径——Diffusion Transformer(DiT)在更少的token上运行。然而,这类自编码器难以训练:更高的压缩率会降低重建质量,恢复质量又需要更多的通道数,而这会减慢DiT的收敛。压缩后的潜变量也与DiT训练时使用的不同,因此预训练DiT必须从头重训或以可观的成本进行适配。压缩DiT训练所用的自编码器看似保持了兼容性,但仅针对重建进行优化仍会将潜变量推离DiT已学到的分布。为此,我们提出GRACE(面向生成的高效视频生成的生成感知潜变量压缩),一个两阶段框架,在保持与预训练DiT兼容的同时压缩预训练视频自编码器。具体而言,我们冻结预训练编码器的基础潜变量,学习一个残差潜变量来捕获更强压缩下丢失的信息,同时在冻结DiT的特征空间中对齐压缩潜变量与预训练潜变量,使自编码器面向生成进行优化。随后通过轻量微调和非对称去噪(基础部分先于残差部分去噪)来适配DiT。GRACE将Wan2.1-I2V-14B的token数量减少8倍,延迟降低11.1倍(480×832×81分辨率),同时在VBench上匹配压缩前管线的生成质量。
原文摘要
Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compre...
自动采集于 2026-10-09
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。