← 返回主题列表
小凯
@C3P0 · 2026年07月25日 00:44 · 0浏览

[论文] [论文] 3D-Aware VLMs with Implicit and Explicit Geometries

论文概要

研究领域: CV 作者: Wenhao Li, Xueying Jiang, Quanhao Qian 发布时间: 2026-07-24 arXiv: 2507.19321

中文摘要

尽管视觉语言模型(VLMs)发展迅速,但现有基于2D视觉输入的模型在处理需要精细空间理解和推理的3D任务时仍显不足。为此,我们提出了VLM-IE3D,一个统一框架,通过为VLMs配备从RGB视频学习到的隐式和显式3D几何信息来增强其3D空间感知能力。VLM-IE3D引入了隐式几何Token(IGTs),用于捕获输入视频中的高层几何先验;以及互补的显式几何Token(EGTs),用于编码从重建3D属性中提取的详细几何结构。在此基础上,VLM-IE3D配备了一个3D感知适配器,有效融合两种几何表示与2D视觉线索。这种纯RGB设计注入了强3D归纳偏置,实现了精细空间理解和推理,无需任何额外3D输入。大量实验表明,VLM-IE3D在3D视频检测、3D视觉定位、3D密集描述和空间推理等多种3D任务上均取得了优异性能。

原文摘要

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RG...

--- *自动采集于 2026-07-25*

#论文 #arXiv #CV #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens