论文概要
研究领域: CV
作者: Agniv Chatterjee, Georgios Pavlakos
发布时间: 2026-08-27
arXiv: 2608.27407
中文摘要
3D人机交互估计(3D HOI)是3D计算机视觉中的基础问题,应用于AR/VR、机器人和具身AI。然而,由于深度模糊、遮挡和物体形状变化,在3D中重建这些交互仍然具有挑战性。现有方法主要关注重投影和接触约束,将参数化人体模型和物体模板拟合到2D图像。本文探索了不同途径。我们提出了MILO,一个利用大型重建模型(LRMs)视觉能力从单张图像恢复详细3D人机交互的框架。我们的关键观察是LRM提供强大的几何支架,保留相对的人-物排列和接近线索。这显著简化了重建过程,将问题重新定义为解释LRM网格:我们将其分割成人体和物体组件,将参数化身体模型拟合到人体部分,并将物体模板对齐到物体部分(如果可用)。MILO在多个基准和交互场景中实现了强大的重建精度并优于现有基线。
原文摘要
Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI. However, reconstructing these interactions in 3D remains challenging due to depth ambiguities, occlusions, and object shape variability. Existing approaches are primarily concerned with reprojection and contact constraints, fitting parametric human models and object templates to 2D images. In this paper, we explore a different avenue. We present MILO, a framework that leverages the visual capabilities of Large Reconstruction Models (LRMs) to recover detailed 3D human-object interactions from a single image. Our key observation is that LRMs provide a powerful geometric scaffold that preserves relative human-object arrangement and proxim...
自动采集于 2026-08-30
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。