论文概要
研究领域: CV
作者: Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
发布时间: 2026-09-14
arXiv: 2609.15980
中文摘要
当视频模型生成物理上不正确的运动时,它是未能学到正确运动,还是学到了但未能使用?我们证明是后者:正确的运动仍然在模型内部可用,仍然可以被用来控制生成的视频。我们在红色质量块缓慢振荡、蓝色质量块快速振荡的视频上训练模型,然后测试一个观测到快速运动的红质量块。即使模型在这种冲突情况下生成了缓慢运动,一个由简单物理变量预测的低维编辑仍能恢复正确的快速运动。我们将这种能力称为"因果可写性"(causal writability)。在固定强度下,我们发现了一个尖锐的深度边界:同一编辑在边界之前可以改变视频,但在边界之后则不能。这个闭合点标志着该写入的承诺。运动信号仍然存在,更强的下游写入可以恢复物理运动,但增益过大会过冲。早期因果可写性可以预测哪些错误会在后续训练中被纠正:这些错误比持续存在的错误在更多网络深度上是可写的。我们在预训练的1.3B视频模型中复现了因果可写性及其尖锐闭合,支持其跨模型规模和训练范式的普遍性。
原文摘要
When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevert...
自动采集于 2026-09-16
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。