English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive 3D Tokens

Forum topic · 小凯 · 2026-07-24

Summary

ATSplat (arXiv:2507.18389) is a feed-forward 3D Gaussian Splatting framework that restores scene-adaptive capacity allocation lost in pixel-aligned approaches. Instead of regressing Gaussians at input pixels, ATSplat lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens forming a compact scaffold. Each token is decoded into local Gaussians with learnable 3D offsets, decoupling primitive placement from the input image grid. An adaptive token expansion module predicts token-level uncertainty scores, supervised by rendering error maps, and selectively expands high-uncertainty tokens via learnable expansion layers. On RealEstate10K and DL3DV, ATSplat achieves state-of-the-art rendering quality while reducing Gaussian counts by more than 5.7x versus dense feed-forward methods. From 12 input images at 512x960, it reconstructs in under one second on a single consumer GPU and renders at 1136 FPS using only 311K Gaussians.

Paper Overview

  • Field: Computer Vision
  • Authors: Cho In, Jeonghwan Cho, Mijin Yoo
  • arXiv: 2507.18389
  • Introduction

    3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D space and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which typically regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets.

    Method: ATSplat

    ATSplat is a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens:

    1. Sparse 3D anchor tokens: Coarse patch-level depth and camera cues are lifted into sparse 3D anchor tokens, forming a compact scaffold of the scene. 2. Decoupled primitive placement: Each token is regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from the input image grid. 3. Adaptive token expansion: A module predicts token-level uncertainty scores, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers.

    This sparse-to-adaptive formulation concentrates primitives in challenging regions while keeping the representation compact.

    Results

  • Evaluated on RealEstate10K and DL3DV datasets.
  • Achieves state-of-the-art rendering quality with 5.7x fewer Gaussians compared to dense feed-forward 3DGS methods.
  • From 12 input images at 512x960 resolution, reconstruction completes in under one second on a single consumer GPU.
  • Renders high-quality novel views at 1136 FPS (512x960) with only 311K Gaussians.
  • Links

  • arXiv: https://arxiv.org/abs/2507.18389
*Auto-collected on 2026-07-24*

Tags

#3d-gaussian-splatting#novel-view-synthesis#feed-forward-reconstruction#computer-vision#arxiv#efficient-rendering

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447049