English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis

Forum topic · 小凯 · 2026-08-19

Summary

SplatGuide is a pose-free novel view synthesis framework that combines feed-forward 3D Gaussian Splatting (3DGS) reconstruction with multi-view diffusion. Unlike prior pipelines that extract only a single signal from the reconstruction, SplatGuide reuses one 3DGS scene in three complementary roles: rendered images provide pixel-aligned geometric conditioning, per-Gaussian source-view indices are rendered into a target-view voting map for occlusion-aware reference selection, and reconstruction tokens supply feature-level guidance via cross-attention. All three signals come from a single reconstruction forward pass. The method achieves state-of-the-art pose-free novel view synthesis on RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360. Notably, on RealEstate10K with a moderate number of input views, SplatGuide surpasses baselines that use ground-truth camera poses. Paper: arXiv 2608.16863, by Yejun Zhang, Zihan Wang, Xu Ji et al.

Overview

Field: Computer Vision Authors: Yejun Zhang, Zihan Wang, Xu Ji et al. (11 authors) Published: 2026-08-17 arXiv: 2608.16863

Abstract

Generating photorealistic novel views from unposed images requires both 3D geometric understanding and the ability to synthesize unseen content. A natural strategy combines feed-forward 3DGS reconstruction with multi-view diffusion. Yet prior pipelines extract at most one signal from the reconstruction — either pixel rendering or learned features — while none exploits per-Gaussian visibility for occlusion-aware reference selection. This information disconnect leaves renderable geometry, visibility cues, and learned features unused.

SplatGuide closes this disconnect by reusing a single 3DGS scene across three complementary roles:

  • Pixel-aligned geometric conditioning: rendered images provide geometry-aware guidance for the diffusion model.
  • Occlusion-aware reference selection: per-Gaussian source-view indices are rendered into a target-view voting map.
  • Feature-level guidance: reconstruction tokens provide learned feature guidance via cross-attention.
  • All three signals are derived from the same single reconstruction forward pass.

    Results

  • State-of-the-art pose-free novel view synthesis on RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360.
  • On RealEstate10K, with a moderate number of input views, SplatGuide outperforms baselines that use ground-truth poses.
--- *Auto-collected on 2026-08-19*

Tags

#novel-view-synthesis#3d-gaussian-splatting#multi-view-diffusion#pose-free#computer-vision#generative-models#3d-reconstruction

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633642