Summary
CityRAG is a video generative model that produces 3D-consistent, navigable simulations of real locations, spatially grounded to actual geography. Unlike standard text-to-video (T2V) or image-to-video (I2V) models that only generate plausible sequences from prompts, CityRAG leverages large corpora of geo-registered data as context to anchor generation to the physical scene, while retaining learned priors for complex motion and appearance changes. A key design choice is its use of temporally unaligned training data, which teaches the model to semantically decouple the underlying scene from transient attributes such as weather, lighting, and dynamic objects. Experiments show that CityRAG generates coherent, minute-long, physically grounded video sequences that maintain consistent weather and lighting conditions across thousands of frames, achieve loop closure, and navigate complex trajectories to reconstruct real-world geography. This capability is essential for downstream applications such as autonomous driving and robotics simulation. The paper (arXiv:2604.19741) is authored by Gene Chou, Charles Herrmann, Kyle Genova, and colleagues from Google and Cornell.
Paper Overview
Research Area: Computer Vision (CV)
Authors: Gene Chou, Charles Herrmann, Kyle Genova, Boyang Deng, Songyou Peng, Bharath Hariharan, Jason Y. Zhang, Noah Snavely, Philipp Henzler
Published: 2026-04-21
arXiv: 2604.19741
Abstract
We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V) prompt. However, the capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for downstream applications including autonomous driving and robotics simulation.
To this end, we present CityRAG, a video generative model that leverages large corpora of geo-registered data as context to ground generation to the physical scene, while maintaining learned priors for complex motion and appearance changes. CityRAG relies on temporally unaligned training data, which teaches the model to semantically decouple the underlying scene from its transient attributes.
Key Results
Experiments demonstrate that CityRAG can generate coherent, minute-long, physically grounded video sequences that:
- Maintain consistent weather and lighting conditions across thousands of frames
- Achieve loop closure
- Navigate complex trajectories to reconstruct real-world geography
This makes the model a promising foundation for applications requiring faithful real-location simulation, such as autonomous driving and robotics.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177618646