论文概要
研究领域: CV
作者: Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni
发布时间: 2026-09-17
arXiv: 2609.20817
中文摘要
从稀疏单目视图建模铰接物体具有挑战性,因为每次观察仅揭示部分几何和运动证据。大多数前馈方法从单一观察推断铰接结构,因此严重依赖学习到的类别级形状先验。我们提出 FAMOS——一种前馈模型,从稀疏、无序的部分点云集合中预测可动部件分割和关节参数。我们的模型对多个观察进行联合推理,天然支持可变数量的输入(包括单视图)。为聚合跨观察的铰接线索,我们引入了多状态铰接 Transformer(Multi-state Articulation Transformer),交替使用状态级和全局注意力。我们进一步提出"观察铰接跨度"目标,监督每个部件在输入观察中表现出的运动范围,鼓励模型利用完整的观察集。为克服现有数据集规模和多样性有限的问题,我们引入了程序化数据生成器,在训练期间合成自标注资产。在 PartNet-Mobility、ACD 和 ArtiCraft-10K 上的实验表明,FAMOS 一致优于前馈和基于优化的基线方法。项目页面:https://kevinqu7.github.io/famos
原文摘要
Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises t...
自动采集于 2026-09-19
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。