English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PhysiFormer: A Diffusion Transformer for Physically Plausible 3D Motion Simulation in World Space

Forum topic · 小凯 · 2026-06-27

Summary

PhysiFormer is a diffusion transformer introduced by Yiming Chen, Yushi Lan, and Andrea Vedaldi (arXiv:2606.27364) that generates physically plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represents objects as 3D meshes expressed directly in world coordinates. Given initial vertex positions and velocities along with the object's material type (rigid or elastic), the model samples future vertex trajectories, producing simulations grounded in physical plausibility. The approach differs from related neural physics methods that rely on ad-hoc latent spaces or explicitly enforce rigidity and causality, instead learning mechanics end-to-end in a shared world-space representation. The paper falls in the computer vision research area and was published on arXiv on 2026-06-27. This forum post shares the paper's abstract and links to its arXiv page for readers interested in generative models, world models, and neural physics simulation.

This forum post shares a new research paper: PhysiFormer: Learning to Simulate Mechanics in World Space.

  • Research area: Computer Vision (CV)
  • Authors: Yiming Chen, Yushi Lan, Andrea Vedaldi
  • Published: 2026-06-27
  • arXiv: 2606.27364
  • Abstract

    We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represents objects as 3D meshes expressed in world coordinates. Given the initial vertex positions and velocities, as well as object material type, rigid or elastic, the model samples future vertex trajectories. While related neural physics approaches build on ad-hoc latent spaces or explicitly enforce rigidity and causality, PhysiFormer learns mechanics directly in a shared world-space representation.

    Key Points

  • Uses a diffusion transformer architecture to model 3D object dynamics.
  • Operates in world coordinates on 3D meshes rather than pixel space, avoiding view-dependent artifacts of video world models.
  • Conditions on initial vertex positions and velocities plus material type (rigid or elastic).
  • Generates future vertex trajectories with physical plausibility.
  • Contrasts with neural physics methods based on ad-hoc latent spaces or explicit rigidity/causality constraints.
*Auto-collected on 2026-06-27.*

Tags

#physiformer#diffusion-transformer#3d-motion#neural-physics#world-models#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208195