Paper Overview
Research Area: CV/Physics
Authors: Tianyu Xu, Shuzhou Yang, Jinbo Xing et al.
Published: 2026-04-30
arXiv: 2604.28169
Abstract
Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a framework that introduces continuous, interpretable, and physically grounded control into video generation.
Our approach integrates three key components:
1. Large-scale physics dataset: over 100K photorealistic simulation videos where friction, restitution, deformation, and forces are systematically varied across diverse scenes. 2. Physics-supervised fine-tuning: a pretrained diffusion model is fine-tuned with a ControlNet conditioned on pixel-aligned physical property maps. 3. VLM-guided reward optimization: a fine-tuned vision-language model evaluates generated videos through targeted physical queries and provides differentiable feedback.
This combination enables generative models to produce physically consistent and controllable outputs through variations of physical attributes — without requiring any simulator or geometric reconstruction at inference time.
Results
On the Physics-IQ benchmark, PhyCo significantly outperforms strong baselines. Human studies confirm clearer, more faithful control over physical attributes. The results demonstrate a scalable path toward physically consistent, controllable generative video models that generalize beyond synthetic training environments.
---
*Auto-collected on 2026-05-02*