English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Goals

Forum topic · 小凯 · 2026-08-15

Summary

PlayWorld is a new benchmark introduced by researchers from the computer vision community to fairly compare interactive video world models. Instead of fixed action-conditioned evaluation—unsuitable because different models need different action sequences to achieve the same outcome—PlayWorld uses multi-modal Agent Players that interact with world models toward specified long-horizon objectives, such as turning 360 degrees to check environmental consistency or walking into water to verify realistic ripple generation. The benchmark provides 171 scenarios, each with a specified objective, and evaluates models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution, plus basic metrics for video quality and controllability. Experiments across nine state-of-the-art world models show that current models remain unreliable on long-horizon interactive objectives, especially in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld. The paper is listed on arXiv as 2608.13552.

Overview

Research area: Computer Vision (CV)

Authors: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

arXiv: 2608.13552

Abstract

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging.

In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison.

The PlayWorld Benchmark

To address this, the authors employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, PlayWorld provides 171 scenarios, each with a specified objective.

Evaluation covers four core dimensions:

  • Geometry consistency
  • Interaction fidelity
  • Out-of-sight evolution
  • Insight evolution
In addition, basic ability metrics for video quality and controllability are incorporated.

Findings

Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution.

Resources

Code and data are available at: https://github.com/kxding/PlayWorld

---

*Automatically collected on 2026-08-15.*

Tags

#world-models#benchmark#computer-vision#video-generation#agent-players#long-horizon-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633498