论文概要
研究领域: CV
作者: Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag
发布时间: 2026-10-01
arXiv: 2610.02197
中文摘要
视频生成模型已取得惊人的视觉保真度,有望成为通用世界模拟器。尽管进展巨大,它们仍无法生成遵循物理定律的视频。在多个物理原理须在同一视频中协同作用的真实场景中,问题更加突出——例如「气球向上飘、锅中蒸汽同时升腾」需要浮力与流体动力学连贯且同时地展开。然而现有方法大多忽略多原理交互,每个视频只关注单一原理。我们提出 HiPhy(层次化物理对齐),一个将视频生成扎根于物理定律的强化学习框架,通过双层目标实现:局部层面强制单个物理原理的时间动力学,全局层面确保整个场景的物理与语义连贯性。为支撑多原理生成,我们构建了 5 万条提示的数据集并引入提示基准 MultiPhyBench,涵盖多样化的共现物理事件。实验表明,HiPhy 显著优于先前方法与基线,在多个基准上大幅改善物理常识与语义对齐,并在竞争性方法退化最剧烈的多物理原理并发场景中取得最大提升。
原文摘要
Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforc...
自动采集于 2026-10-04
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。