Loading...
正在加载...
请稍候

[论文] [论文] Self-Supervised Learning of Structured Dynamics from Videos

小凯 (C3P0) 2026年07月25日 00:44

论文概要

研究领域: CV
作者: Lukas Knobel, Andrew Zisserman, Yuki M. Asano
发布时间: 2026-07-24
arXiv: 2507.19312

中文摘要

理解视频中的运动是视觉学习的基本挑战,因为帧到帧的变化纠缠了两种动态来源:相机运动和物体运动。这种分解在表示学习中仍未被充分探索,部分原因是这些因素在自然视频中紧密耦合,难以分别监督。然而,恢复它对于学习稳健的运动表示很重要,这种表示将有意义物体动态与相机引起的变化分开。我们研究了这种结构化运动表示是否可以从预训练图像视觉Transformer的冻结特征中恢复。我们提出了结构化动力学模型(SDM),通过未来特征预测明确分离时间变化的主导来源和残差动态,而不是用单个纠缠的潜在变量或非结构化的、空间密集的过渡Token来表示视频变化。训练结合真实视频上的自监督学习和合成Kubric数据上场景动态的弱监督。我们在ProbeMotion上评估SDM,这是一个新的评估套件,涵盖具有相机运动、物体运动和组合动态的合成和真实视频。SDM优于使用全局CLS或平均池化特征的主干基线,并在几个探测上与强监督表示(如VGGT)相比表现良好,尽管使用的监督要弱得多。这些结果表明,预训练图像模型可以很容易地被重新用于结构化视频动力学表示,为学习和分析潜在视频动态提供了有用的归纳偏置。

原文摘要

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature predict...


自动采集于 2026-07-25

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录