Summary
VLM-IE3D is a unified framework that enhances the 3D spatial awareness of vision-language models (VLMs) by equipping them with both implicit and explicit 3D geometries learned purely from RGB videos. The paper, authored by Wenhao Li, Xueying Jiang, and Quanhao Qian (arXiv:2507.19321), addresses the weakness of existing 2D-input VLMs on tasks requiring fine-grained spatial understanding. The framework introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, and complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. A 3D-aware adapter fuses these two geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases without requiring any additional 3D inputs such as depth sensors or LiDAR. Experiments show VLM-IE3D achieves strong performance across diverse 3D tasks, including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning.
Paper Overview
- Field: Computer Vision
- Authors: Wenhao Li, Xueying Jiang, Quanhao Qian
- arXiv: 2507.19321
Abstract
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, the authors present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos.
Key Contributions
- Implicit Geometry Tokens (IGTs): capture high-level geometric priors from input videos.
- Explicit Geometry Tokens (EGTs): encode detailed geometric structures from reconstructed 3D attributes, complementing the IGTs.
- 3D-aware adapter: effectively fuses the two types of geometric representations with 2D visual cues.
This RGB-only design injects strong 3D inductive biases, enabling fine-grained spatial understanding and reasoning without requiring any additional 3D inputs.
Results
Extensive experiments demonstrate that VLM-IE3D achieves strong performance across a variety of 3D tasks, including:
- 3D video detection
- 3D visual grounding
- 3D dense captioning
- Spatial reasoning
---
*Auto-collected on 2026-07-25*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178447080