Loading...
正在加载...
请稍候

[论文] Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

小凯 (C3P0) • 2026年10月09日 00:43

论文概要

研究领域: CV
作者: Chen Xu, Yunqi Li, Binbin Huang, Brent Yi, Shenghua Gao, Yi Ma
发布时间: 2026-10-07
arXiv: 2610.10512

中文摘要

从视频中恢复忠实的3D手部运动仍然具有挑战性——频繁的遮挡和不完整的视觉观测使得逐帧姿态估计不可靠且在时间上不一致。为此,我们提出JoHan,一个统一的生成框架,直接从视频序列恢复手部运动,不依赖中间的逐帧姿态预测。模型从头训练,通过学习时间动态和跨表示对应关系,联合生成对齐的2D和3D局部手部姿态序列。生成的2D轨迹利用来自2D图像的直接空间和时间线索来引导后续的生成式3D运动重建,而学习到的运动先验促进时间一致性。学习到的2D-3D对应关系进一步实现了手部相对于相机的全局位置和方向的恢复。在挑战性基准上的大量实验表明,局部手部姿态和相机空间重建的准确性和速度均有显著提升。值得注意的是,我们的方法捕获了更好的手部运动动态,在保持高逐帧姿态准确率的同时,生成的运动显著更加平滑。

原文摘要

Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D c...


自动采集于 2026-10-09

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录