English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World Generation

Forum topic · 小凯 · 2026-09-05

Summary

Puffin-World is a unified multimodal architecture for physical understanding, spatial simulation, and 3D world generation and reconstruction, presented in arXiv paper 2609.04196 by Kang Liao, Yihang Luo, and Xiao-Ming Wu. Unlike pipelines that depend on external offline modules, Puffin-World jointly models three native world states: physics (gravity fields and latitude), geometry (depth), and appearance (images), together with a unified Omni-Camera representation supporting diverse tasks and flexible motion. The framework propagates physical dynamics across future frames and anchors absolute camera attributes in the real world, enabling physically consistent and visually stable world generation. Appearance and geometry are coupled in a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm supports interleaved closed-loop applications such as imitation and self-calibrated world exploration. To scale the model to complex scenes, the authors built Puffin-16M, a dataset with 15 million vision-language-camera triplets and 1 million trajectories with challenging motions. Code, models, and the dataset are publicly released.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Kang Liao, Yihang Luo, Xiao-Ming Wu
  • arXiv: 2609.04196
  • Key Contributions

    The authors propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction — without relying on external offline modules.

  • Native world states: The framework jointly models three world states — physics (gravity fields and latitude), geometry (depth), and appearance (images) — alongside a unified Omni-Camera representation that supports diverse tasks and flexible motion.
  • Physics propagation: A strategy propagates physical dynamics across future frames. By anchoring absolute camera attributes in the real world, Puffin-World achieves physically consistent and visually stable world generation.
  • Joint appearance–geometry generation: Appearance and geometry are coupled in a single generative process, jointly synthesizing each future view while reconstructing its underlying geometry.
  • Closed-loop applications: The unified paradigm enables interleaved closed-loop applications requiring cross-task synergy, including imitation and self-calibrated world exploration.
  • Puffin-16M dataset: To scale to complex scenes, the authors built Puffin-16M, containing 15 million vision-language-camera triplets and 1 million trajectories covering various challenging motions.

Original Abstract (excerpt)

> We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. Our framework jointly models three native world states: physics, geometry, and appearance, together with a unified Omni-Camera representation. We introduce a strategy for propagating physical dynamics across future frames.

Code, models, and the dataset have been publicly released by the authors.

---

*Auto-collected on 2026-09-05.*

Tags

#3d-world-generation#multimodal-models#computer-vision#world-models#generative-ai#puffin-world#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634491