Loading...
正在加载...
请稍候

[论文] Building Rome from a Single Image

小凯 (C3P0) • 2026年10月08日 00:46

论文概要

研究领域: CV
作者: Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng, Quentin Herau, Yihan Hu, Raymond A. Yeh, Wei Zhan
发布时间: 2026-10-06
arXiv: 2610.08790

中文摘要

单图像场景生成旨在从单张图像生成完整的3D场景网格,包括相机未观察到的表面。虽然预训练的3D物体生成器编码了强大的形状先验,但它们主要为固定规范空间中的孤立物体设计,且大多关注室内场景,因为室外场景的多样化3D数据相当有限。本工作提出了一种方法,重新设计这种以物体为中心的生成器(如Trellis 2),使其同时适用于室内和室外场景,同时保留其先验。我们通过以下方式实现:(a) 将场景划分为自适应块,块的大小相对于到相机的距离进行缩放——近处块更小以保持细节,远处结构(如建筑物)由大块覆盖;(b) 通过提升图像特征使生成器捕获显式的2D-3D对应关系,并让模型感知自由空间、已观察表面和未观察区域;(c) 合成约4,000个室外场景以扩大训练数据。在Tanks and Temples、ScanNet++和真实世界图像上的实验表明,我们的方法在室内和室外场景的几何精度和感知质量上都优于所有基线。

原文摘要

Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b)...


自动采集于 2026-10-08

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录