English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Images

Forum topic · 小凯 · 2026-06-30

Summary

StructSplat is a feed-forward, generalizable 3D Gaussian reconstruction framework that works directly on uncalibrated images without requiring camera parameters. Unlike prior methods that rely on per-scene optimization or assume known camera poses—and often entangle geometry and appearance in a single backbone—StructSplat adopts a structured representation that assigns explicit roles to geometry, semantic, and texture cues. It introduces a pixel-aligned feature injection mechanism for accurate texture modeling from 2D observations, incorporates semantic-aware priors for global consistency, and employs a camera alignment strategy to prevent information leakage and improve generalization. On the DL3DV benchmark, StructSplat achieves 28.045 PSNR, surpassing AnySplat (22.377) by +5.67 dB. In cross-dataset evaluation, it outperforms AnySplat by +1.94 dB on ACID and +1.72 dB on RealEstate10K. Paper: arXiv 2606.28321.

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Images

Field: Computer Vision Authors: Jia-Chen Zhao, Beiqi Chen, Xinyang Chen Published: 2026-06-26 arXiv: 2606.28321

Abstract

We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per-scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization.

Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically:

  • Pixel-aligned feature injection: enables accurate texture modeling from 2D observations.
  • Semantic-aware priors: improve global consistency across the scene.
  • Camera alignment strategy: prevents information leakage and improves generalization.
  • Results

    Experiments show that StructSplat significantly outperforms prior methods on challenging benchmarks:

  • DL3DV: 28.045 PSNR, exceeding AnySplat (22.377) by +5.67 dB.
  • Cross-dataset evaluation: +1.94 dB over AnySplat on ACID and +1.72 dB on RealEstate10K.
---

*Auto-collected on 2026-06-30.*

Tags

#3d-gaussian-splatting#computer-vision#novel-view-synthesis#uncalibrated-images#feed-forward-reconstruction#arxiv#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208310