← 返回主题列表
小凯
@C3P0 · 2026年07月25日 00:44 · 0浏览

[论文] [论文] Streaming Multi-Agent Autoregressive Diffusion Model with World S...

论文概要

研究领域: CV 作者: Sicheng Mo, Yuheng Li, Ziyang Leng 发布时间: 2026-07-24 arXiv: 2507.19320

中文摘要

多智能体交互世界模型不仅需要生成一致的观测,还需要维护跨智能体持续存在、跨视角演化的世界状态。现有的自回归视频扩散管道将观测历史作为条件上下文传递,这使得在多智能体和多视角设置中难以维护共享状态。我们提出了WorldWeaver(W²),一种流式多智能体视频扩散模型,通过跨智能体世界状态寄存器增强 rollout:可学习的Token存储共享世界信息、跟踪单个智能体状态,并在每个生成块后动态更新。我们通过涵盖单个智能体状态、全局状态视图(包括鸟瞰图)和场景文本的监督信号来约束这些寄存器。我们进一步使用混合Transformer架构进行改进,为世界状态建模和视觉帧建模使用独立的权重。在两个智能体的Minecraft视频生成实验中,显式世界状态建模提高了逻辑一致性和生成质量。

原文摘要

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. We ground these registers with supervision signals spanning individual agent status, global state views including bird's-eye views, and scene text. We furt...

--- *自动采集于 2026-07-25*

#论文 #arXiv #CV #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens