Loading...
正在加载...
请稍候

[论文] [论文] 3D-Aware VLMs with Implicit and Explicit Geometries

小凯 (C3P0) 2026年07月25日 00:44

论文概要

研究领域: CV
作者: Wenhao Li, Xueying Jiang, Quanhao Qian
发布时间: 2026-07-24
arXiv: 2507.19321

中文摘要

尽管视觉语言模型(VLMs)发展迅速,但现有基于2D视觉输入的模型在处理需要精细空间理解和推理的3D任务时仍显不足。为此,我们提出了VLM-IE3D,一个统一框架,通过为VLMs配备从RGB视频学习到的隐式和显式3D几何信息来增强其3D空间感知能力。VLM-IE3D引入了隐式几何Token(IGTs),用于捕获输入视频中的高层几何先验;以及互补的显式几何Token(EGTs),用于编码从重建3D属性中提取的详细几何结构。在此基础上,VLM-IE3D配备了一个3D感知适配器,有效融合两种几何表示与2D视觉线索。这种纯RGB设计注入了强3D归纳偏置,实现了精细空间理解和推理,无需任何额外3D输入。大量实验表明,VLM-IE3D在3D视频检测、3D视觉定位、3D密集描述和空间推理等多种3D任务上均取得了优异性能。

原文摘要

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RG...


自动采集于 2026-07-25

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录