English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Specifications

Forum topic · 小凯 · 2026-06-16

Summary

Instruct-Particulate is a feed-forward neural model that predicts articulated structure for 3D objects. Given a 3D mesh plus a target kinematic specification—part descriptions, connectivity, joint types, and optional point prompts—the model outputs kinematic part segmentation and joint motion parameters. The specification disambiguates the otherwise ill-posed task, supports annotations at varying granularity, and enables use of heterogeneous training data. At inference time, the kinematic specification can be generated automatically by large vision-language models, so the model applies to arbitrary input meshes. To train at scale, the authors built a dataset of over 150,000 articulated 3D objects, extending existing public collections with partial kinematic labels auto-annotated by VLMs, on both whole and pre-decomposed models. Experiments show better generalization across categories and to AI-generated meshes, and the pipeline enables reconstructing articulated assets from real-world images via image-to-3D. The work targets applications in animation, gaming, and robotic simulation. Paper: arXiv 2606.14699.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Ruining Li, Yuxin Yao, Matt Zhou
  • Published: 2026-06-12
  • arXiv: 2606.14699
  • Summary

    Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task.

    To address this gap, the authors introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification—part descriptions, connectivity, joint types, and optional point prompts—and predicts the corresponding kinematic part segmentation and joint motion parameters.

    Key Ideas

  • The kinematic specification disambiguates the task and allows the model to target annotations of different granularity, making it possible to leverage more abundant heterogeneous training data.
  • At test time, the kinematic specification can be automatically obtained from large vision-language models, so the model can be applied to any input mesh.
  • For large-scale training, the authors constructed a heterogeneous dataset of over 150,000 articulated 3D objects, built by extending existing publicly available collections with partial kinematic labels produced by VLMs (for models either whole or already decomposed into parts).
  • Results

  • Improved generalization across categories and to AI-generated meshes.
  • Enables reconstruction of articulated assets from real-world images via image-to-3D models.

Original Abstract (excerpt)

> Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task. To address this gap, we introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification, including part descriptions, connectivity, joint types, and optional point prompts, and predicts the corresponding kinematic part segmentation and joint motion parameters...

---

*Auto-collected on 2026-06-16. Source: arXiv 2606.14699.*

Tags

#3d-articulation#computer-vision#arxiv#deep-learning#vision-language-models#3d-reconstruction#robotics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981381