English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLM-IE3D: 3D-Aware Vision-Language Models with Implicit and Explicit Geometries

Forum topic · 小凯 · 2026-07-27

Summary

VLM-IE3D is a unified framework that enhances vision-language models (VLMs) with 3D spatial awareness by incorporating both implicit and explicit 3D geometries learned purely from RGB videos. Most existing VLMs rely on 2D visual inputs and struggle with 3D tasks requiring fine-grained spatial understanding and reasoning. The framework introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, alongside complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. A 3D-aware adapter fuses these geometric representations with 2D visual cues. Because it requires only RGB input, VLM-IE3D injects strong 3D inductive biases without extra 3D data. Experiments show superior and consistent performance across 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. Paper: arXiv 2507.21747.

Paper Overview

Research Area: Computer Vision (CV) Authors: Wenhao Li, Xueying Jiang, Quanhao Qian Published: 2025-07-27 arXiv: 2507.21747

Abstract

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, the authors present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos.

Key components:

  • Implicit Geometry Tokens (IGTs) — capture high-level geometric priors from input videos.
  • Explicit Geometry Tokens (EGTs) — encode detailed geometric structures extracted from reconstructed 3D attributes, complementing IGTs.
  • 3D-aware adapter — effectively fuses the two types of geometric representations with 2D visual cues.
  • This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D input. Extensive experiments demonstrate that VLM-IE3D achieves superior and consistent performance across diverse 3D tasks, including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning.

    Key Takeaways

  • Addresses the gap between 2D-trained VLMs and 3D spatial reasoning tasks.
  • Combines implicit and explicit geometry via a single token-based framework and a dedicated 3D-aware adapter.
  • Requires only RGB video — no point clouds, depth maps, or other 3D sensors.
  • Validated on multiple 3D benchmarks with strong, stable results.
Source: arXiv:2507.21747

Tags

#vision-language-models#3d-understanding#computer-vision#spatial-reasoning#geometric-representations#arxiv#research-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503705