Paper Overview
Research Area: Computer Vision (CV) Authors: Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi Published: 2026-07-06 arXiv: 2607.05392
Abstract (English)
This paper introduces SynCity 3000, a framework for generating globally consistent 3D scenes with fine-grained layout control. The approach leverages the ability of current image-to-3D generators to produce complex 3D assets from a single image, and extends this capability to the scene level by adapting the generator into a convolutional operator.
Key Ideas
- Scene-scale generation via convolutional adaptation: A pretrained image-to-3D generator is repurposed as a convolutional operator, so that it can be applied beyond single objects to entire scene images.
- Synthetic data engine: To address the scarcity of 3D scene training data, the model is fine-tuned on scene-level data produced by a new synthetic data engine.
- Axonometric image conditioning: The convolutional generator is applied to an axonometric image of a whole scene generated from a user prompt, enabling 3D scenes of arbitrary scale and complexity.
Results
Across diverse prompts and layouts, SynCity 3000 generates large, coherent, and detailed scenes, overcoming shortcomings of previous 3D scene generation methods.
--- *Automatically collected on 2026-07-06*