Loading...
正在加载...
请稍候

[论文] Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Repre...

小凯 (C3P0) 2026年09月07日 01:15

论文概要

研究领域: CV
作者: Denis M. Akola, David F. Fouhey
发布时间: 2026-09-06
arXiv: 2509.04284

中文摘要

3D基础模型(3DFM)如VGGT最近通过前馈Transformer预测丰富的统一表示,推动了3D视觉的边界。这些模型学到的场景表示使其在多个3D视觉任务上表现优异。本文研究如何利用其内部表示从新视角推断场景中的3D信息。我们的假设是:为了解决3D重建任务,这些模型需要学习一种包含大量关于3D场景通用知识的表示。在展示了可以从3DFM内部表示解码隐藏表面后,我们提出了一种名为Z3D的方法,通过对3DFM表示进行潜在扩散来估计未见过视角的点图。我们证明Z3D能够在多个数据集上为新视角预测逼真的深度图。

原文摘要

3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predi...


自动采集于 2026-09-07

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录