English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Forum topic · 小凯 · 2026-09-19

Summary

FAMOS is a feed-forward model for predicting movable-part segmentation and joint parameters of articulated 3D objects from a sparse, unordered set of partial point clouds. Unlike most feed-forward methods that rely on a single observation and heavy category-level shape priors, FAMOS jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. The authors introduce a Multi-state Articulation Transformer that alternates state-wise and global attention to aggregate articulation cues across observations, plus an observed articulation span objective that supervises each part's motion range visible in the inputs, encouraging full use of the observation set. To overcome limited dataset scale and diversity, they propose a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K show FAMOS consistently outperforms both feed-forward and optimization-based baselines. Paper: arXiv:2609.20817.

Overview

Field: Computer Vision (CV) Authors: Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni Published: 2026-09-17 arXiv: 2609.20817

Abstract (translated)

Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to exploit the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K show that FAMOS consistently outperforms both feed-forward and optimization-based baselines.

Links

  • Paper: https://arxiv.org/abs/2609.20817
  • Project page: https://kevinqu7.github.io/famos

Tags

#3d-vision#articulated-objects#feed-forward-models#transformers#point-clouds#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634976