English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PhysStream: Physics-Grounded Streaming Video Generation with Structured Scene Memory

Forum topic · 小凯 · 2026-09-17

Summary

PhysStream is an autoregressive image-to-video generation model that enables interactive, physics-grounded control of dynamic scenes during streaming generation. Developed by researchers including Chuhao Chen, Peter Wonka, Sergey Tulyakov, and Lingjie Liu, the method introduces a structured scene memory consisting of positional maps and object tracking maps derived online from previously generated frames. Instead of pixel-space position signals, PhysStream uses sparse velocity-increment signals encoding physical quantities, allowing the model to learn underlying dynamics. Training proceeds in two stages: first fine-tuning a bidirectional model with motion-control conditioning, then training a causal autoregressive model augmented with structured scene memory to improve physical consistency. PhysStream supports mid-stream interactive control of multi-object tabletop rigid-body scenarios, a capability previous methods lacked. On synthetic benchmarks it reduces the Fréchet Video Motion Distance (FVMD) by 33% and trajectory error by 12% compared with the strongest baselines, and human evaluators preferred its results in over 85% of real-world comparisons. Project page: https://czzzzh.github.io/PhysStream

Paper Overview

Field: Computer Vision (CV) Authors: Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu arXiv: 2609.17521

Summary

Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that dictate object positions rather than physical dynamics.

To address these limitations, the authors propose PhysStream, an autoregressive model for physics-grounded image-to-video synthesis that features:

  • Structured scene memory: positional maps and object tracking maps derived online from previously generated frames.
  • Fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.
  • The model is trained in two stages:

    1. Fine-tuning a bidirectional model with motion-control conditioning. 2. Training a causal autoregressive model augmented with structured scene memory to further improve physical consistency.

    Results

  • Enables mid-stream interactive control of multi-object tabletop rigid-body scenarios — a capability not supported by prior methods.
  • Reduces motion distribution distance (FVMD) by 33% on synthetic benchmarks.
  • Reduces trajectory error by 12% (both improvements over the strongest baseline).
  • Preferred by human evaluators in over 85% of real-world comparisons.
  • Links

  • arXiv: https://arxiv.org/abs/2609.17521
  • Project page: https://czzzzh.github.io/PhysStream
---

*Auto-collected on 2026-09-17.*

Tags

#video-generation#physics-simulation#autoregressive-models#computer-vision#interactive-control#scene-memory#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634899