English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PhysiFormer: Learning to Simulate Mechanics in World Space

Forum topic · 小凯 · 2026-06-27

Summary

PhysiFormer is a diffusion transformer model for generating physically plausible 3D object motion, presented in a paper by Yiming Chen, Yushi Lan, and Andrea Vedaldi (arXiv 2606.27364). Unlike video world models that operate in view-dependent pixel space, PhysiFormer represents objects directly as 3D meshes in world coordinates. Given initial vertex positions and velocities together with an object's material type—rigid or elastic—the model samples future vertex trajectories. Instead of building on ad-hoc latent spaces or explicitly enforcing rigidity and causality like prior neural physics approaches, PhysiFormer learns physical behavior from data in a shared world space, allowing unified treatment of different materials and long-horizon physically consistent simulation.

Overview

Research area: Computer Vision Authors: Yiming Chen, Yushi Lan, Andrea Vedaldi Published: 2026-06-27 arXiv: 2606.27364

Abstract

We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represents objects as 3D meshes expressed in world coordinates. Given the initial vertex positions and velocities, as well as object material type, rigid or elastic, the model samples future vertex trajectories. While related neural physics approaches build on ad-hoc latent spaces or explicitly enforce rigidity and causality, PhysiFormer learns physics directly from data in world space.

*(Source abstract appears truncated in the original post.)*

Key Points

  • PhysiFormer is a diffusion transformer designed for physically plausible 3D object motion simulation.
  • It operates in world coordinates on 3D meshes, in contrast to video world models working in view-dependent pixel space.
  • Inputs include initial vertex positions and velocities and the object's material type (rigid or elastic).
  • The model learns physical behavior from data rather than explicitly enforcing rigidity or causality constraints.

Tags

#physics-simulation#diffusion-transformer#3d-motion#computer-vision#world-models#neural-physics#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208181