English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PAR3D: A Unified Part-Aware 3D-MLLM for Scene Understanding

Forum topic · 小凯 · 2026-06-07

Summary

PAR3D is a unified part-aware 3D multimodal large language model (3D-MLLM) framework designed to understand, reason about, and ground both objects and their component parts in 3D scenes. While existing 3D-MLLMs handle visual question answering, captioning, and referring segmentation, they remain largely object-centric and struggle with fine-grained part structures needed for embodied interaction. To address this, the authors introduce ScenePart, a synthetic 3D scene dataset with part-level annotations and language instructions for training and evaluation. The framework features part-aware 3D representation learning that enriches visual features with part-level semantics, plus hierarchical segmentation query generation that grounds part-level targets through hierarchical object-part queries. Extensive experiments show significant improvements in part-level question answering and referring segmentation, while maintaining strong performance on object-level vision-language tasks. arXiv: 2606.06485.

Paper Overview

Research Area: Computer Vision Authors: Shaohui Dai, Yansong Qu, You Shen Published: 2026-06-04 arXiv: 2606.06485

Abstract

Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning, and referring segmentation. However, existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures that are essential for embodied interaction with 3D environments.

Contributions

  • PAR3D framework: A unified part-aware 3D-MLLM that enables models to understand, reason about, and ground both objects and their parts in 3D scenes.
  • ScenePart dataset: A synthetic 3D scene dataset with part-level annotations and language instructions, supporting training and evaluation of part-aware 3D scene understanding.
  • Part-aware 3D representation learning: Enriches 3D visual representations with fine-grained part-level semantics.
  • Hierarchical segmentation query generation: Grounds part-level targets via hierarchical object-part queries.

Results

Extensive experiments demonstrate that PAR3D significantly improves part-level question answering and referring segmentation, while also achieving strong performance on object-level vision-language tasks.

---

*Auto-collected on 2026-06-07*

Tags

#3d-mlm#scene-understanding#computer-vision#multimodal-llm#part-aware-representation#referring-segmentation#embodied-ai#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980917