English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Images

Forum topic · 小凯 · 2026-06-30

Summary

StructSplat is a feed-forward, generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters, as presented by Jia-Chen Zhao, Beiqi Chen, and Xinyang Chen in arXiv paper 2606.28321. Unlike prior methods that rely on per-scene optimization or assume known camera poses, StructSplat uses a structured representation that assigns explicit roles to geometry, semantic, and texture cues during reconstruction. It introduces a pixel-aligned feature injection mechanism for accurate texture modeling from 2D observations, semantic-aware priors to improve global consistency, and a camera alignment strategy to prevent information leakage and improve generalization. Experiments show significant improvements over prior methods: on DL3DV it reaches 28.045 PSNR, exceeding AnySplat (22.377) by +5.67 dB; in cross-dataset evaluation it outperforms AnySplat by +1.94 dB on ACID and +1.72 dB on RealEstate10K.

Paper Overview

Field: Computer Vision (CV) Authors: Jia-Chen Zhao, Beiqi Chen, Xinyang Chen Published: 2026-06-26 arXiv: 2606.28321

Abstract

We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per-scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization.

Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically:

  • Pixel-aligned feature injection mechanism — enables accurate texture modeling from 2D observations
  • Semantic-aware priors — improve global consistency
  • Camera alignment strategy — prevents information leakage and improves generalization
  • Results

    Experiments show that our method significantly outperforms previous approaches on challenging benchmarks:

  • On DL3DV, our method reaches 28.045 PSNR, exceeding AnySplat (22.377) by +5.67 dB
  • In cross-dataset evaluation, our method outperforms AnySplat by +1.94 dB on ACID and +1.72 dB on RealEstate10K
--- *Auto-collected on 2026-06-30*

Tags

#3d-gaussian-splatting#computer-vision#novel-view-synthesis#3d-reconstruction#uncalibrated-images#feed-forward#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208299