[论文] PhysStream: Streaming Physics-Grounded Video Generation with Structure...

研究领域: CV 作者: Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu 发布时间: 2026-09-15 arXiv: 2609.17521

论文概要

研究领域: CV 作者: Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu 发布时间: 2026-09-15 arXiv: 2609.17521

中文摘要

视频生成的交互式控制正从粗略提示走向对动态场景细粒度、具有物理意义的操控。然而现有的可控方法要么需要在生成开始前就确定完整的控制时间表,要么使用直接规定物体位置的像素空间信号,而非物理动力学。为解决这些局限,我们提出 PhysStream——一种基于物理的图像到视频合成自回归模型,它引入了结构化场景记忆(从已生成帧中在线导出的位置图和物体跟踪图),并支持通过稀疏速度增量信号进行细粒度运动控制,这些信号编码物理量,让模型学习底层动力学。我们分两阶段训练模型:首先用运动控制条件对双向模型进行微调,然后在增加结构化场景记忆的基础上训练因果自回归模型,进一步提升物理一致性。PhysStream 实现了对多物体桌面刚体场景的中途交互式控制——这是此前方法不支持的能力——在合成基准上将运动分布距离(FVMD)降低了33%,轨迹误差降低了12%(均优于最强基线),并在超过85%的真实世界比较中获得了人类评估者的偏好。详情请访问我们的网站:https://czzzzh.github.io/PhysStream

原文摘要

Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that dictate object positions rather than physical dynamics. To address these limitations, we propose PhysStream, an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory---positional maps and object tracking maps derived online from previously generated frames---and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics. We train our model in two stages: a bidirectio...


*自动采集于 2026-09-17*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens