Loading...
正在加载...
请稍候

[论文] [论文] ROWBench: Do Video Models Render What the Program Specifies?

小凯 (C3P0) • 2026年10月03日 00:42

论文概要

研究领域: CV
作者: Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
发布时间: 2026-10-01
arXiv: 2610.02205

中文摘要

可编程世界模型将可执行动力学与视觉生成分离,为下一代游戏引擎提供了有前景的基础。然而它们对显式规则和交互的视觉遵循程度仍未得到充分评估。现有基准评估视觉质量、可控性和指令或物理遵循性,但很少测试对细粒度程序指定世界事件的保真度。我们提出 PROWBench,包含 170 个程序化构建的剧集和 600 个代理视频,覆盖多样场景和交互。PROWBench 记录实体状态和时间戳事件(包括视野外事件),作为可重放的世界记录,从中渲染同步视图和代理表示。这使得生成的视频可以与程序执行的可观察结果进行核对。可扩展框架构建场景、控制行为,并可将每个相机视图渲染为不同表示(如粗粒度 3D 和边界框)。基准覆盖第一和第三人称视角,部分剧集提供同步多视图观察。基于这些记录,PROWBench 评估实体控制、长时程记忆,并通过两个基于 VLM 的指标——逻辑-渲染对齐和交互成功率——评估对规定时间线和时间戳引擎记录事件的视觉实现程度。

原文摘要

Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be che...


自动采集于 2026-10-03

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录