English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldWeaver (W²): Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Forum topic · 小凯 · 2026-07-25

Summary

This paper introduces WorldWeaver (W²), a streaming multi-agent video diffusion model for interactive world modeling, presented by Sicheng Mo, Yuheng Li, and Ziyang Leng (arXiv 2507.19320). Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes it difficult to maintain shared world states across multiple agents and viewpoints. WorldWeaver augments the rollout process with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. These registers are grounded via supervision signals covering individual agent status, global state views (including bird's-eye views), and scene text. The model further employs a hybrid Transformer architecture with separate weights for world state modeling and visual frame modeling. Experiments on Minecraft video generation with two agents show that explicit world state modeling improves both logical consistency and generation quality.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Sicheng Mo, Yuheng Li, Ziyang Leng
  • Published: 2026-07-24
  • arXiv: 2507.19320
  • Abstract

    Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings.

    The authors present WorldWeaver (W²), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk.

    Key Contributions

  • Cross-agent world state registers: learnable tokens that store shared world information and track individual agent status across the generation process.
  • Supervised grounding: registers are constrained with supervision signals spanning individual agent status, global state views (including bird's-eye views), and scene text.
  • Hybrid Transformer architecture: separate weights are used for world state modeling and visual frame modeling.
  • Results: experiments on Minecraft video generation with two agents demonstrate that explicit world state modeling improves logical consistency and generation quality.

Original Abstract (excerpt)

> Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk...

*Source: arXiv:2507.19320*

Tags

#world-models#video-diffusion#multi-agent#autoregressive#transformer#minecraft#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447081