[论文] Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
研究领域: NLP 作者: Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc…
论文概要
研究领域: NLP 作者: Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong 发布时间: 2026-09-29 arXiv: 2609.38177
中文摘要
从多视角图像推理3D世界对多模态大语言模型(MLLM)而言仍是一个根本性挑战。虽然现代MLLM能有效处理单图像输入,但它们难以将跨视角证据整合为连贯的3D理解。越来越多的工作试图通过向MLLM注入3D感知来弥合这一差距—— either 通过增强细粒度像素级跨视图对应,或通过融合3D几何基础模型的特征——但与人类推理之间仍存在显著差距。本文重新审视人类的空间推理过程:人类并非依赖细粒度几何线索,而是粗略地识别跨视角的共同物体、推断视角间的相对几何关系,并组装场景的粗略3D布局。受此启发,我们提出Imagine3D-LLM——一种学习组装类似的紧凑3D场景表示并以此为基础生成回答的MLLM。具体而言,我们在图像token后附加少量可学习的摘要token,将其解码为紧凑的3D高斯溅射表示,并通过光度重建损失进行监督,同时与标准的下一token预测目标联合训练。值得注意的是,虽然仅摘要token接收直接的重建监督,但该目标也在LLM底层图像特征中诱导了更强的跨帧对应性,表明学习重建将3D感知信号传播到了整个模型。结果表明,Imagine3D-LLM在多个空间推理和3D理解基准上一致优于先前方法,暗示'想象场景'可能比'被告知像素级几何'更有效。
原文摘要
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of...
*自动采集于 2026-10-01*
#论文 #arXiv #NLP #小凯