论文概要
研究领域: CV
作者: Wenhao Li, Xueying Jiang, Quanhao Qian
发布时间: 2026-07-24
arXiv: 2507.19321
中文摘要
尽管视觉语言模型(VLMs)发展迅速,但现有基于2D视觉输入的模型在处理需要精细空间理解和推理的3D任务时仍显不足。为此,我们提出了VLM-IE3D,一个统一框架,通过为VLMs配备从RGB视频学习到的隐式和显式3D几何信息来增强其3D空间感知能力。VLM-IE3D引入了隐式几何Token(IGTs),用于捕获输入视频中的高层几何先验;以及互补的显式几何Token(EGTs),用于编码从重建3D属性中提取的详细几何结构。在此基础上,VLM-IE3D配备了一个3D感知适配器,有效融合两种几何表示与2D视觉线索。这种纯RGB设计注入了强3D归纳偏置,实现了精细空间理解和推理,无需任何额外3D输入。大量实验表明,VLM-IE3D在3D视频检测、3D视觉定位、3D密集描述和空间推理等多种3D任务上均取得了优异性能。
原文摘要
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RG...
自动采集于 2026-07-25
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。