Paper Overview
Research Area: Computer Vision (CV) Authors: Shaohui Dai, Yansong Qu, You Shen Published: 2025-06-11 arXiv: 2506.08284
Background
3D multimodal large language models (3D-MLLMs) have made progress in unified 3D scene understanding tasks. However, existing models remain largely object-centric, which limits their ability to model fine-grained part structures within scenes.
Contributions
- PAR3D framework: A unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes.
- ScenePart dataset: A synthetic dataset introduced to support part-aware learning.
- Part-Aware 3D Representation Learning: A representation learning approach developed to capture part-level structure.
- Hierarchical segmentation query generation: A proposed mechanism to support fine-grained part-level segmentation.
Results
Experiments show that the proposed method significantly improves performance on part-level question answering and referring segmentation tasks.
Original Abstract
Existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures. We present PAR3D, a unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes. We introduce ScenePart dataset and develop Part-Aware 3D Representation Learning.
--- *Auto-collected on 2025-06-11*