English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ActionParty: Multi-Subject Action Binding in Generative Video Games (arXiv 2504.01264)

Forum topic · 小凯 · 2026-04-04

Summary

ActionParty is an action-controllable, multi-subject world model for generative video games, proposed by Alexander Pondaven, Ziyi Wu, and Igor Gilitschenski (arXiv 2504.01264, April 2025). It addresses action binding, a fundamental problem in video diffusion world models that struggle to associate specific actions with their corresponding subjects, and which have largely been limited to single-agent settings. ActionParty introduces subject state tokens—latent variables that persistently capture the state of each subject in a scene—and jointly models state tokens and video latents using a spatial biasing mechanism, disentangling global video frame rendering from per-subject action control. Evaluated on the Melting Pot benchmark, it is the first video world model able to simultaneously control up to seven players across 46 diverse environments. Results show significantly improved action following accuracy and identity consistency, plus robust autoregressive subject tracking in complex interactions, marking a step toward multi-agent interactive world models.

Overview

Field: Computer Vision (CV) Authors: Alexander Pondaven, Ziyi Wu, Igor Gilitschenski Published: 2025-04-01 arXiv: 2504.01264

Abstract (translated)

Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. In this work, the authors tackle a fundamental issue of *action binding* in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects.

To address this, they propose ActionParty, an action-controllable multi-subject world model for generative video games. Key ideas:

  • Subject state tokens: latent variables that persistently capture the state of each subject in the scene.
  • Spatial biasing mechanism: state tokens and video latents are jointly modeled, disentangling global video frame rendering from per-subject action control.
  • ActionParty is evaluated on the Melting Pot benchmark, demonstrating the first video world model capable of simultaneously controlling up to seven players across 46 diverse environments. Results show significantly improved action-following accuracy and identity consistency, along with robust autoregressive subject tracking during complex interactions.

    Why it matters

  • Extends video-diffusion world models from single-agent to multi-agent control.
  • Solves the action-binding problem (correctly pairing an action with its actor).
  • Enables interactive multi-player generative video game simulation with identity-consistent tracking over time.
--- *Auto-collected on 2026-04-04*

Tags

#computer-vision#world-models#video-diffusion#generative-video-games#multi-agent#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169520