[论文] CtrlCache: Accelerating Interactive Video World Models with Control-Aw...
研究领域: CV 作者: Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang 发布时间: 2026-10-06 arXiv: 2610.08777
论文概要
研究领域: CV 作者: Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang 发布时间: 2026-10-06 arXiv: 2610.08777
中文摘要
交互式视频世界模型需要高效生成每个视频块,同时忠实地响应用户控制。许多系统使用分块自回归生成与少步去噪,但每个块仍需要多次昂贵的去噪迭代。无需训练的缓存可以减少这种成本,但现有策略主要基于模型内部的去噪动态做重用决策,并未明确考虑控制转换。实际上,交互式生成明确暴露了一个它们未使用的信号:块的控制在去噪之前到达,因此从它们派生的调度不消耗前向传播。我们分析了不同控制方案下的相邻块,发现结构相似性在动作变化时下降,而低频结构比高频细节更持久。基于这些观察,我们提出CtrlCache,一个无需训练的控制感知缓存框架,使计算适应当前的控制序列。具体而言,动作感知的调度和刷新策略检测跨块和块内的动作变化,并将每个块标记为初始、过渡、转向或稳定状态。在一个选定的内部去噪步骤,初始和过渡块保留完整计算,而转向和稳定块重用同一块中最近完整计算步骤的transformer残差。为了在稳定交互期间利用低频结构的持久性,我们进一步引入频率混合的历史先验引导,无需额外DiT前向传播即可融入来自前一个干净latent的补充信息。在Matrix-Game 2.0和LingBot-World v1/v2上的评估中,CtrlCache在无需模型重训练的情况下实现了1.21倍到1.41倍的DiT骨干加速,同时在所有三个模型上提高了WBench总分。
原文摘要
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequ...
*自动采集于 2026-10-08*
#论文 #arXiv #CV #小凯