English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLM-IE3D: 3D-Aware VLMs with Implicit and Explicit Geometries

Forum topic · 小凯 · 2026-07-26

Summary

VLM-IE3D is a unified framework that enhances the 3D spatial awareness of vision-language models (VLMs) by equipping them with both implicit and explicit 3D geometries learned purely from RGB videos. The method introduces Implicit Geometry Tokens (IGTs), which capture high-level geometric priors from input videos, and complementary Explicit Geometry Tokens (EGTs), which encode detailed geometric structures from reconstructed 3D attributes. A 3D-aware adapter fuses these two geometric representations with 2D visual cues, injecting strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D inputs such as depth sensors or point clouds. Experiments show consistent state-of-the-art results across 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning benchmarks. The paper is authored by Wenhao Li, Xueying Jiang, and Quanhao Qian (arXiv:2507.20491), with code and models publicly available on GitHub.

Paper Overview

Field: Computer Vision Authors: Wenhao Li, Xueying Jiang, Quanhao Qian Published: 2026-07-25 arXiv: 2507.20491

Introduction

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, the authors present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos.

Key Components

  • Implicit Geometry Tokens (IGTs): capture high-level geometric priors from input videos.
  • Explicit Geometry Tokens (EGTs): encode detailed geometric structures derived from reconstructed 3D attributes, complementing the IGTs.
  • 3D-aware adapter: effectively fuses the two types of geometric representations with 2D visual cues.
  • This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning, without requiring any additional 3D inputs (e.g., depth sensors or point clouds).

    Results

    Extensive experiments show that VLM-IE3D achieves consistently strong performance across a range of 3D tasks:

  • 3D video detection
  • 3D visual grounding
  • 3D dense captioning
  • Spatial reasoning
  • Resources

  • Paper: https://arxiv.org/abs/2507.20491
  • Code and models: https://github.com/Vegetebird/VLM-IE3D
--- *Auto-collected on 2026-07-26*

Tags

#vlm#3d-awareness#computer-vision#rgb-video#spatial-reasoning#geometry-tokens#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447116