[论文] Building Rome from a Single Image

研究领域: CV 作者: Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng, Quentin Herau, Yihan Hu, Raymond A. Yeh, Wei Zhan 发布时间: 2026-10-06 arXiv: 2610.08790

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng, Quentin Herau, Yihan Hu, Raymond A. Yeh, Wei Zhan 发布时间: 2026-10-06 arXiv: 2610.08790

中文摘要

单图像场景生成旨在从单张图像生成完整的3D场景网格,包括相机未观察到的表面。虽然预训练的3D物体生成器编码了强大的形状先验,但它们主要为固定规范空间中的孤立物体设计,且大多关注室内场景,因为室外场景的多样化3D数据相当有限。本工作提出了一种方法,重新设计这种以物体为中心的生成器(如Trellis 2),使其同时适用于室内和室外场景,同时保留其先验。我们通过以下方式实现:(a) 将场景划分为自适应块,块的大小相对于到相机的距离进行缩放——近处块更小以保持细节,远处结构(如建筑物)由大块覆盖;(b) 通过提升图像特征使生成器捕获显式的2D-3D对应关系,并让模型感知自由空间、已观察表面和未观察区域;(c) 合成约4,000个室外场景以扩大训练数据。在Tanks and Temples、ScanNet++和真实世界图像上的实验表明,我们的方法在室内和室外场景的几何精度和感知质量上都优于所有基线。

原文摘要

Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b)...


*自动采集于 2026-10-08*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens