[论文] FlowHMR: Physically Plausible Motion Capture from Video

研究领域: CV 作者: Zhanke Wang, Chengfeng Zhao, Qing Shuai, Jingzhong Lin, Heng Li, Zeyu Ling, Yuxin Wen, Jing Li, Di Kang, Chunchao Guo, Linchao Bao 发布时间: 2026-10-0…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Zhanke Wang, Chengfeng Zhao, Qing Shuai, Jingzhong Lin, Heng Li, Zeyu Ling, Yuxin Wen, Jing Li, Di Kang, Chunchao Guo, Linchao Bao 发布时间: 2026-10-02 arXiv: 2610.03691

中文摘要

我们提出 FlowHMR——一个从单目视频中恢复物理上合理的全局三维人体运动的框架。先前的学习方法通常直接从视频回归人体运动,并用几何监督训练网络。然而,从单目视频恢复人体运动在深度上存在固有的歧义,直接回归倾向于坍缩到平均解。此外,恢复的运动不保证物理合理,基于物理的跟踪经常失败。为应对这些挑战,我们将视频动作捕捉建模为视频条件运动生成问题,首先为该任务预训练一个流匹配模型。给定输入视频,预训练模型生成多样的运动候选,但并非所有候选都忠实于视频或可被物理跟踪。因此,我们使用 Group Relative Policy Optimization(GRPO)和两个奖励对模型进行后训练:忠实性奖励鼓励与输入视频的一致性;跟踪奖励偏好能被物理控制器成功跟踪的运动。这两个奖励共同将模型的输出偏好转移,使后训练模型在忠实于输入视频的同时产生更物理合理的运动。我们还引入了 Wild-4K——一个约 4000 个网络视频的大规模多样数据集,用于评估野外人体运动恢复。Wild-4K 上的定性和定量实验表明,我们的方法在整体运动忠实性上优于最先进水平,物理跟踪成功率为 82.47%,而最强基线 GVHMR 为 62.82%。

原文摘要

We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are...


*自动采集于 2026-10-06*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens