[论文] CtrlCache: Accelerating Interactive Video World Models with Control-Aw...

研究领域: CV 作者: Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang 发布时间: 2026-10-06 arXiv: 2610.08777

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang 发布时间: 2026-10-06 arXiv: 2610.08777

中文摘要

交互式视频世界模型需要高效生成每个视频块,同时忠实地响应用户控制。许多系统使用分块自回归生成与少步去噪,但每个块仍需要多次昂贵的去噪迭代。无需训练的缓存可以减少这种成本,但现有策略主要基于模型内部的去噪动态做重用决策,并未明确考虑控制转换。实际上,交互式生成明确暴露了一个它们未使用的信号:块的控制在去噪之前到达,因此从它们派生的调度不消耗前向传播。我们分析了不同控制方案下的相邻块,发现结构相似性在动作变化时下降,而低频结构比高频细节更持久。基于这些观察,我们提出CtrlCache,一个无需训练的控制感知缓存框架,使计算适应当前的控制序列。具体而言,动作感知的调度和刷新策略检测跨块和块内的动作变化,并将每个块标记为初始、过渡、转向或稳定状态。在一个选定的内部去噪步骤,初始和过渡块保留完整计算,而转向和稳定块重用同一块中最近完整计算步骤的transformer残差。为了在稳定交互期间利用低频结构的持久性,我们进一步引入频率混合的历史先验引导,无需额外DiT前向传播即可融入来自前一个干净latent的补充信息。在Matrix-Game 2.0和LingBot-World v1/v2上的评估中,CtrlCache在无需模型重训练的情况下实现了1.21倍到1.41倍的DiT骨干加速,同时在所有三个模型上提高了WBench总分。

原文摘要

Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequ...


*自动采集于 2026-10-08*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens