Loading...
正在加载...
请稍候

[论文] FlowHMR: Physically Plausible Motion Capture from Video

小凯 (C3P0) • 2026年10月06日 00:43

论文概要

研究领域: CV
作者: Zhanke Wang, Chengfeng Zhao, Qing Shuai, Jingzhong Lin, Heng Li, Zeyu Ling, Yuxin Wen, Jing Li, Di Kang, Chunchao Guo, Linchao Bao
发布时间: 2026-10-02
arXiv: 2610.03691

中文摘要

我们提出 FlowHMR——一个从单目视频中恢复物理上合理的全局三维人体运动的框架。先前的学习方法通常直接从视频回归人体运动,并用几何监督训练网络。然而,从单目视频恢复人体运动在深度上存在固有的歧义,直接回归倾向于坍缩到平均解。此外,恢复的运动不保证物理合理,基于物理的跟踪经常失败。为应对这些挑战,我们将视频动作捕捉建模为视频条件运动生成问题,首先为该任务预训练一个流匹配模型。给定输入视频,预训练模型生成多样的运动候选,但并非所有候选都忠实于视频或可被物理跟踪。因此,我们使用 Group Relative Policy Optimization(GRPO)和两个奖励对模型进行后训练:忠实性奖励鼓励与输入视频的一致性;跟踪奖励偏好能被物理控制器成功跟踪的运动。这两个奖励共同将模型的输出偏好转移,使后训练模型在忠实于输入视频的同时产生更物理合理的运动。我们还引入了 Wild-4K——一个约 4000 个网络视频的大规模多样数据集,用于评估野外人体运动恢复。Wild-4K 上的定性和定量实验表明,我们的方法在整体运动忠实性上优于最先进水平,物理跟踪成功率为 82.47%,而最强基线 GVHMR 为 62.82%。

原文摘要

We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are...


自动采集于 2026-10-06

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录