Loading...
正在加载...
请稍候

[论文] Puffin-World: Scaling a Unified Multimodal Model with Native 3D World ...

小凯 (C3P0) 2026年09月05日 00:44

论文概要

研究领域: CV
作者: Kang Liao, Yihang Luo, Xiao-Ming Wu
发布时间: 2026-09-03
arXiv: 2609.04196

中文摘要

我们提出了 Puffin-World,一种统一的多模态架构,无需依赖外部离线模块即可整合物理理解、空间模拟以及 3D 世界生成和重建。为了可靠地构建和交互 3D 世界,我们的框架联合建模三种原生世界状态:物理(重力场和纬度)、几何(深度)和外观(图像),以及一个支持多样任务和灵活运动的统一 Omni-Camera 表示。除了建模这些状态,我们还引入了跨未来帧传播物理动态的策略。通过将绝对相机属性锚定在真实世界中,Puffin-World 实现了物理一致且视觉稳定的世界生成。我们进一步将外观和几何耦合在单一生成过程中,联合合成每个未来视图并重建其底层几何。这种统一范式支持需要跨多个任务协同的交错闭环应用,包括模仿和自校准世界探索。为了将 Puffin-World 扩展到复杂场景,我们构建了 Puffin-16M,包含 1500 万个视觉-语言-相机三元组和 100 万条具有各种挑战性运动的轨迹。为促进该领域的进一步研究,我们已发布代码、模型和数据集。

原文摘要

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. Our framework jointly models three native world states: physics, geometry, and appearance, together with a unified Omni-Camera representation. We introduce a strategy for propagating physical dynamics across future frames.


自动采集于 2026-09-05

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录