Paper Overview
Research Area: Computer Vision Authors: Shaohui Dai, Yansong Qu, You Shen Published: 2026-06-04 arXiv: 2606.06485
Abstract
Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning, and referring segmentation. However, existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures that are essential for embodied interaction with 3D environments.
Contributions
- PAR3D framework: A unified part-aware 3D-MLLM that enables models to understand, reason about, and ground both objects and their parts in 3D scenes.
- ScenePart dataset: A synthetic 3D scene dataset with part-level annotations and language instructions, supporting training and evaluation of part-aware 3D scene understanding.
- Part-aware 3D representation learning: Enriches 3D visual representations with fine-grained part-level semantics.
- Hierarchical segmentation query generation: Grounds part-level targets via hierarchical object-part queries.
Results
Extensive experiments demonstrate that PAR3D significantly improves part-level question answering and referring segmentation, while also achieving strong performance on object-level vision-language tasks.
---
*Auto-collected on 2026-06-07*