Loading...
正在加载...
请稍候

[论文] Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

小凯 (C3P0) 2026年08月15日 00:47

论文概要

研究领域: CV
作者: Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao
发布时间: 2026-08-13
arXiv: 2608.13546

中文摘要

交互式世界模型必须支持持久记忆、响应式交互和长程生成,但这些要求对模型提出了相互冲突的需求。在去噪器上下文或键值缓存中维护历史会产生不断增长的成本,迫使在会话长度和保留记忆之间进行权衡,而低延迟交互依赖于少步生成,其能力受限于教师模型。Evoke通过外化持久世界状态并为长程交互生成重新设计教师来解决这两个限制。场景几何维护在外部的、相机索引的世界状态库中,仅检索与视图相关的信息,使去噪器上下文在会话增长时保持有界。我们不将教师视为固定生成器,而是为其设计长程监督:其稀疏注意力结合块级分组、选定远端帧的检索和线性注意力全局状态,在内存和计算上实现线性增长,同时实现对长程的监督。这种监督暴露了在短窗口内局部合理的内容漂移,而每块条件化实现整个序列中的提示变化和事件控制。在自强制rollout下应用的30秒分布匹配目标将两种能力转移到一个不使用无分类器引导的三步学生模型,提高对长期漂移的抵抗能力同时保留响应式条件化。凭借有界上下文和循环外部记忆,Evoke支持开放式、持续演化的生成;在单张H200上以384×640分辨率,每1.5秒块在2.11秒内生成。作为三步世界模型,Evoke在WBench上达到最先进性能,同时在VBench-Long和VBench-2.0上保持竞争力。

原文摘要

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term漂移while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at 384×640, each 1.5 s chunk is generated in 2.11 s. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.


自动采集于 2026-08-15

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录