English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

Forum topic · 小凯 · 2026-04-11

Summary

Scal3R (arXiv:2504.07077) tackles large-scale 3D scene reconstruction from long video sequences. Feed-forward reconstruction models regress 3D geometry directly from RGB images but degrade on long sequences due to limited memory and poor global context. Inspired by how humans use global scene understanding to guide local perception, Scal3R introduces a neural global context representation that compresses and retains long-range scene information. It is realized via lightweight sub-networks adapted at test time through self-supervised objectives, substantially increasing memory capacity with minimal computational overhead. Experiments on KITTI Odometry and Oxford Spires benchmarks show leading pose accuracy and state-of-the-art 3D reconstruction accuracy on ultra-large scenes while maintaining efficiency.

Paper Overview

  • Field: AI / 3D Reconstruction
  • Authors: Tao Xie, Peishan Yang, Yudong Jin
  • Published: 2025-04-10
  • arXiv: 2504.07077

Abstract

This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or geometric constraints. However, these methods often struggle to maintain reconstruction accuracy and consistency over long sequences due to limited memory capacity and the inability to effectively capture global contextual cues. In contrast, humans can naturally exploit the global understanding of the scene to inform local perception.

Motivated by this, the authors propose a novel neural global context representation that efficiently compresses and retains long-range scene information, enabling the model to leverage extensive contextual cues for enhanced reconstruction accuracy and consistency. The context representation is realized through a set of lightweight neural sub-networks that are rapidly adapted during test time via self-supervised objectives, which substantially increases memory capacity without incurring significant computational overhead.

Results

Experiments on multiple large-scale benchmarks, including the KITTI Odometry and Oxford Spires datasets, demonstrate the effectiveness of the approach in handling ultra-large scenes, achieving leading pose accuracy and state-of-the-art 3D reconstruction accuracy while maintaining efficiency.

--- *Auto-collected on 2025-04-11*

Tags

#3d-reconstruction#test-time-training#computer-vision#deep-learning#pose-estimation#feed-forward-models#kitti#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169737