[论文] FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

研究领域: CV 作者: Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni 发布时间: 2026-09-17 arXiv: 2609.20817

论文概要

研究领域: CV 作者: Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni 发布时间: 2026-09-17 arXiv: 2609.20817

中文摘要

从稀疏单目视图建模铰接物体具有挑战性,因为每次观察仅揭示部分几何和运动证据。大多数前馈方法从单一观察推断铰接结构,因此严重依赖学习到的类别级形状先验。我们提出 FAMOS——一种前馈模型,从稀疏、无序的部分点云集合中预测可动部件分割和关节参数。我们的模型对多个观察进行联合推理,天然支持可变数量的输入(包括单视图)。为聚合跨观察的铰接线索,我们引入了多状态铰接 Transformer(Multi-state Articulation Transformer),交替使用状态级和全局注意力。我们进一步提出"观察铰接跨度"目标,监督每个部件在输入观察中表现出的运动范围,鼓励模型利用完整的观察集。为克服现有数据集规模和多样性有限的问题,我们引入了程序化数据生成器,在训练期间合成自标注资产。在 PartNet-Mobility、ACD 和 ArtiCraft-10K 上的实验表明,FAMOS 一致优于前馈和基于优化的基线方法。项目页面:https://kevinqu7.github.io/famos

原文摘要

Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises t...


*自动采集于 2026-09-19*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens