English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Guiding Image-to-3D Generation with Test-Time Partial Observations

Forum topic · 小凯 · 2026-09-11

Summary

Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by available observations, limiting use in applications requiring geometric fidelity. This paper by Jerred Chen, Simon Weber, and Ronald Clark (arXiv:2609.10531) introduces a training-free framework that incorporates partial geometric observations available at test time into pretrained image-to-3D generative models, without retraining or finetuning. The method guides generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence. Applied to SAM 3D and its multi-view extension, the approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. The results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative prior without modifying the underlying model. This work was shared on zhichai.net on 2026-09-11.

Paper Overview

  • Field: Computer Vision
  • Authors: Jerred Chen, Simon Weber, Ronald Clark
  • Published: 2026-09-09
  • arXiv: 2609.10531
  • Abstract

    Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. In many real-world settings, however, partial geometric observations of the object may be available at test time.

    The authors introduce a training-free framework for incorporating such evidence into pretrained image-to-3D generative models without retraining or finetuning. To do this, they guide generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence.

    Applied to SAM 3D and its multi-view extension, the approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. The results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative prior without modifying the underlying model.

    Key Contributions

  • Training-free test-time guidance framework for image-to-3D generation
  • Ray-consistent observation likelihood over occupancy representations, using both surface and free-space evidence
  • Significant gains in geometric fidelity and visual quality on SAM 3D and its multi-view extension
--- *Auto-collected on 2026-09-11.*

Tags

#image-to-3d#3d-generation#computer-vision#test-time-guidance#sam-3d#occupancy#generative-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634710