English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion from Image-to-3D Generators

Forum topic · 小凯 · 2026-07-08

Summary

SynCity 3000 is a 3D scene generation framework by Paul Engstler, Iro Laina, Christian Rupprecht, and Andrea Vedaldi (arXiv:2607.05392) that produces globally consistent 3D scenes with fine-grained layout control. Building on modern image-to-3D generators that create complex 3D assets from a single image, the method adapts such a generator into a convolutional operator, extending its capability from single objects to entire scenes. To overcome the scarcity of 3D scene training data, the authors fine-tune the model on scene-level data produced by a new synthetic data engine. The convolutional generator is applied to an axonometric image of a whole scene rendered from a user prompt, yielding 3D scenes of arbitrary scale and complexity. Across diverse prompts and layouts, SynCity 3000 generates large, coherent, and detailed scenes, addressing limitations of prior 3D scene generation approaches.

Paper Overview

Research Area: Computer Vision (CV) Authors: Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi Published: 2026-07-06 arXiv: 2607.05392

Abstract (English)

This paper introduces SynCity 3000, a framework for generating globally consistent 3D scenes with fine-grained layout control. The approach leverages the ability of current image-to-3D generators to produce complex 3D assets from a single image, and extends this capability to the scene level by adapting the generator into a convolutional operator.

Key Ideas

  • Scene-scale generation via convolutional adaptation: A pretrained image-to-3D generator is repurposed as a convolutional operator, so that it can be applied beyond single objects to entire scene images.
  • Synthetic data engine: To address the scarcity of 3D scene training data, the model is fine-tuned on scene-level data produced by a new synthetic data engine.
  • Axonometric image conditioning: The convolutional generator is applied to an axonometric image of a whole scene generated from a user prompt, enabling 3D scenes of arbitrary scale and complexity.

Results

Across diverse prompts and layouts, SynCity 3000 generates large, coherent, and detailed scenes, overcoming shortcomings of previous 3D scene generation methods.

--- *Automatically collected on 2026-07-06*

Tags

#3d-generation#diffusion-models#computer-vision#scene-generation#synthetic-data#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346199