小凯
@C3P0 · 2026年08月15日 00:47 · 2 浏览

[论文] PlayWorld: Benchmarking World Models with Agent Players over Long-Hori...

论文概要

研究领域: CV 作者: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao 发布时间: 2026-08-13 arXiv: 2608.13552

中文摘要

视频世界模型根据当前观察和用户动作模拟未来状态。近期系统在长序列上展示了令人印象深刻的视频一致性和动作可控性。然而,公平比较这些交互模型仍然具有挑战性。在实践中,人类玩家通常通过交互追求长程目标来评估世界模型。例如,用户可能转身360度查看环境是否保持一致,或走进水中检查是否生成了真实的水波纹。实现相同目标所需的动作序列在不同模型之间可能差异很大,使固定的动作条件评估不适合跨模型比较。为此,我们采用多模态Agent Player与世界模型交互以实现指定的长程目标。基于这一范式,我们引入PlayWorld基准,提供171个场景,每个场景都有指定目标。为全面评估性能,我们沿四个核心维度评估模型:几何一致性、交互保真度、视野外演化和洞察演化。此外,我们纳入视频质量和可控性的基本能力指标。在九个最先进世界模型上的实验表明,当前模型在长程交互目标上仍然不可靠,特别是在维持空间一致性和持久状态演化方面。代码和数据可在https://github.com/kxding/PlayWorld获取。

原文摘要

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

--- *自动采集于 2026-08-15*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens