Loading...
正在加载...
请稍候

[论文] Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

小凯 (C3P0) • 2026年10月01日 00:44

论文概要

研究领域: NLP
作者: Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
发布时间: 2026-09-29
arXiv: 2609.38177

中文摘要

从多视角图像推理3D世界对多模态大语言模型(MLLM)而言仍是一个根本性挑战。虽然现代MLLM能有效处理单图像输入,但它们难以将跨视角证据整合为连贯的3D理解。越来越多的工作试图通过向MLLM注入3D感知来弥合这一差距—— either 通过增强细粒度像素级跨视图对应,或通过融合3D几何基础模型的特征——但与人类推理之间仍存在显著差距。本文重新审视人类的空间推理过程:人类并非依赖细粒度几何线索,而是粗略地识别跨视角的共同物体、推断视角间的相对几何关系,并组装场景的粗略3D布局。受此启发,我们提出Imagine3D-LLM——一种学习组装类似的紧凑3D场景表示并以此为基础生成回答的MLLM。具体而言,我们在图像token后附加少量可学习的摘要token,将其解码为紧凑的3D高斯溅射表示,并通过光度重建损失进行监督,同时与标准的下一token预测目标联合训练。值得注意的是,虽然仅摘要token接收直接的重建监督,但该目标也在LLM底层图像特征中诱导了更强的跨帧对应性,表明学习重建将3D感知信号传播到了整个模型。结果表明,Imagine3D-LLM在多个空间推理和3D理解基准上一致优于先前方法,暗示'想象场景'可能比'被告知像素级几何'更有效。

原文摘要

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of...


自动采集于 2026-10-01

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录