English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PhyCo: Learning Controllable Physical Priors for Generative Motion

Forum topic · 小凯 · 2026-05-02

Summary

PhyCo is a new framework that brings continuous, interpretable, physically grounded control into video diffusion generation, addressing common physical inconsistencies such as object drift, unrealistic collision rebounds, and material responses that fail to match underlying physical properties. The method combines three components: (i) a large-scale dataset of over 100K photorealistic simulation videos with systematically varied friction, restitution, deformation, and forces across diverse scenes; (ii) physics-supervised fine-tuning of a pretrained diffusion model using a ControlNet conditioned on pixel-aligned physical property maps; and (iii) VLM-guided reward optimization, where a fine-tuned vision-language model evaluates generated videos via targeted physical queries and provides differentiable feedback. This enables physically consistent and controllable outputs without any simulator or geometric reconstruction at inference time. On the Physics-IQ benchmark, PhyCo significantly outperforms strong baselines, and human studies confirm clearer, more faithful control over physical attributes, suggesting a scalable path toward physics-consistent controllable video generation that generalizes beyond synthetic training environments.

Paper Overview

Research Area: CV/Physics

Authors: Tianyu Xu, Shuzhou Yang, Jinbo Xing et al.

Published: 2026-04-30

arXiv: 2604.28169

Abstract

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a framework that introduces continuous, interpretable, and physically grounded control into video generation.

Our approach integrates three key components:

1. Large-scale physics dataset: over 100K photorealistic simulation videos where friction, restitution, deformation, and forces are systematically varied across diverse scenes. 2. Physics-supervised fine-tuning: a pretrained diffusion model is fine-tuned with a ControlNet conditioned on pixel-aligned physical property maps. 3. VLM-guided reward optimization: a fine-tuned vision-language model evaluates generated videos through targeted physical queries and provides differentiable feedback.

This combination enables generative models to produce physically consistent and controllable outputs through variations of physical attributes — without requiring any simulator or geometric reconstruction at inference time.

Results

On the Physics-IQ benchmark, PhyCo significantly outperforms strong baselines. Human studies confirm clearer, more faithful control over physical attributes. The results demonstrate a scalable path toward physically consistent, controllable generative video models that generalize beyond synthetic training environments.

---

*Auto-collected on 2026-05-02*

Tags

#video-generation#diffusion-models#physical-consistency#controlnet#vlm#physics-simulation#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619040