English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PAR3D: A Unified Part-Aware 3D Multimodal Large Language Model

Forum topic · 小凯 · 2026-06-06

Summary

PAR3D is a unified part-aware 3D multimodal large language model (3D-MLLM) framework proposed to overcome the object-centric limitations of existing 3D-MLLMs. While current models can understand 3D scenes at the object level, they struggle to model fine-grained part structures. PAR3D enables models to understand, reason about, and ground both objects and their constituent parts in 3D scenes. The authors introduce the ScenePart synthetic dataset for part-level supervision, develop Part-Aware 3D Representation Learning, and propose hierarchical segmentation query generation to support part-level grounding. Experiments demonstrate that PAR3D significantly improves performance on part-level question answering and referring segmentation tasks in 3D scenes. The work was posted on arXiv (2506.08284) on June 11, 2025, in the computer vision field.

Paper Overview

Research Area: Computer Vision (CV) Authors: Shaohui Dai, Yansong Qu, You Shen Published: 2025-06-11 arXiv: 2506.08284

Background

3D multimodal large language models (3D-MLLMs) have made progress in unified 3D scene understanding tasks. However, existing models remain largely object-centric, which limits their ability to model fine-grained part structures within scenes.

Contributions

  • PAR3D framework: A unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes.
  • ScenePart dataset: A synthetic dataset introduced to support part-aware learning.
  • Part-Aware 3D Representation Learning: A representation learning approach developed to capture part-level structure.
  • Hierarchical segmentation query generation: A proposed mechanism to support fine-grained part-level segmentation.

Results

Experiments show that the proposed method significantly improves performance on part-level question answering and referring segmentation tasks.

Original Abstract

Existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures. We present PAR3D, a unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes. We introduce ScenePart dataset and develop Part-Aware 3D Representation Learning.

--- *Auto-collected on 2025-06-11*

Tags

#3d-mllm#multimodal-llm#computer-vision#3d-scene-understanding#part-aware-learning#referring-segmentation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980878