[论文] 4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reco...

研究领域: CV 作者: Shiqi Li, Sean Cho, Yijie Li, Fengzhi Guo, Bowen Wen, Cheng Zhang 发布时间: 2026-10-06 arXiv: 2610.08782

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Shiqi Li, Sean Cho, Yijie Li, Fengzhi Guo, Bowen Wen, Cheng Zhang 发布时间: 2026-10-06 arXiv: 2610.08782

中文摘要

现有的4D手-物体重建方法通常依赖于昂贵的逐序列优化,而生成方法通常从随机噪声合成交互,可能导致不稳定的交互预测。我们引入4D-HOF,这是一个前馈框架,从视觉基础模型产生的粗略但有信息量的估计中重建4D手-物体交互。具体而言,我们学习了一个条件flow matching模型,将基础模型导出的手-物体状态传输到交互流形上,使模型能够以前馈方式纠正平移、旋转和对齐中的错误。我们生成式公式的一个关键优势是自然实现了传输过程中的测试时引导——不是在建模后应用单独的事后优化,而是直接使用物理交互约束和观察到的2D证据来引导演化的生成状态,使重建作为生成过程本身的一部分得到细化。通过在多样化数据集上训练生成模型,4D-HOF能够稳健地泛化到具有挑战性的真实场景。在域外基准测试上的实验表明,4D-HOF达到了最先进的性能,产生了更稳定和准确的4D手-物体重建。

原文摘要

Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a sepa...


*自动采集于 2026-10-08*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens