Loading...
正在加载...
请稍候

[论文] Guiding Image-to-3D Generation with Test-Time Partial Observations

小凯 (C3P0) 2026年09月11日 00:44

论文概要

研究领域: CV
作者: Jerred Chen, Simon Weber, Ronald Clark
发布时间: 2026-09-09
arXiv: 2609.10531

中文摘要

图像到3D模型可以从单张RGB图像生成视觉上引人注目的3D资产,但其几何形状通常仅由可用观测松散约束,限制了它们在需要几何保真度的应用中的使用。然而,在许多现实场景中,测试时可能可以获得物体的部分几何观测。本文提出一种无需训练的框架,将此类证据整合到预训练的图像到3D生成模型中,无需重新训练或微调。为此,作者使用在模型占据表示上定义的射线一致观测似然来引导生成,结合表面占据和自由空间证据。应用于SAM 3D及其多视图扩展,该方法在不同可观测水平上大幅改善几何保真度和视觉质量。结果表明,预训练图像到3D模型可以通过显式测试时引导有效整合部分几何观测,补充其学习的生成先验而无需修改底层模型。

原文摘要

Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. In many real-world settings, however, partial geometric observations of the object may be available at test time. We introduce a training-free framework for incorporating such evidence into pretrained image-to-3D generative models without retraining or finetuning. To do this, we guide generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence. Applied to SAM 3D and its multi-view extension, our approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. Our results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative prior without modifying the underlying model.


自动采集于 2026-09-11

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录