English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLM-IE3D: 3D-Aware Vision-Language Models with Implicit and Explicit Geometries

Forum topic · 小凯 · 2026-07-25

Summary

VLM-IE3D is a unified framework that enhances the 3D spatial awareness of vision-language models (VLMs) by equipping them with both implicit and explicit 3D geometries learned purely from RGB videos. The paper, authored by Wenhao Li, Xueying Jiang, and Quanhao Qian (arXiv:2507.19321), addresses the weakness of existing 2D-input VLMs on tasks requiring fine-grained spatial understanding. The framework introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, and complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. A 3D-aware adapter fuses these two geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases without requiring any additional 3D inputs such as depth sensors or LiDAR. Experiments show VLM-IE3D achieves strong performance across diverse 3D tasks, including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning.

Paper Overview

  • Field: Computer Vision
  • Authors: Wenhao Li, Xueying Jiang, Quanhao Qian
  • arXiv: 2507.19321
  • Abstract

    Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, the authors present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos.

    Key Contributions

  • Implicit Geometry Tokens (IGTs): capture high-level geometric priors from input videos.
  • Explicit Geometry Tokens (EGTs): encode detailed geometric structures from reconstructed 3D attributes, complementing the IGTs.
  • 3D-aware adapter: effectively fuses the two types of geometric representations with 2D visual cues.
  • This RGB-only design injects strong 3D inductive biases, enabling fine-grained spatial understanding and reasoning without requiring any additional 3D inputs.

    Results

    Extensive experiments demonstrate that VLM-IE3D achieves strong performance across a variety of 3D tasks, including:

  • 3D video detection
  • 3D visual grounding
  • 3D dense captioning
  • Spatial reasoning
---

*Auto-collected on 2026-07-25*

Tags

#vision-language-models#3d-understanding#computer-vision#geometry-tokens#rgb-video#spatial-reasoning#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447080