English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GLD: Repurposing Geometric Foundation Models for Multi-View Diffusion

Forum topic · 小凯 · 2026-03-25

Summary

GLD (Geometric Latent Diffusion) is a framework that repurposes the geometrically consistent feature space of geometric foundation models as the latent space for multi-view diffusion, targeting novel view synthesis (NVS). Existing NVS approaches typically operate in view-independent VAE latent spaces, which lack cross-view geometric consistency. The authors show that geometric foundation model features support high-fidelity RGB reconstruction while encoding strong cross-view geometric correspondences, making them a well-suited latent space for NVS. Experiments demonstrate that GLD outperforms VAE and RAE baselines on both 2D image quality and 3D consistency metrics, and accelerates training by more than 4.4x compared to VAE latent spaces. Notably, although GLD's diffusion model is trained from scratch without generative pretraining, it remains competitive with state-of-the-art methods that leverage large-scale text-to-image pretraining. Authors: Wooseok Jang, Seonghu Jeon, Jisang Han, Jinhyeok Choi, Minkyung Kwon, Seungryong Kim, Saining Xie, Sainan Liu. arXiv: 2603.22275.

Paper Overview

Field: Computer Vision Authors: Wooseok Jang, Seonghu Jeon, Jisang Han, Jinhyeok Choi, Minkyung Kwon, Seungryong Kim, Saining Xie, Sainan Liu Published: 2026-03-23 arXiv: 2603.22275

Summary

While recent advances in generative latent spaces have driven substantial progress in single-image generation, the optimal latent space for novel view synthesis (NVS) remains largely unexplored. In particular, NVS requires geometrically consistent generation across viewpoints, but existing approaches typically operate in a view-independent VAE latent space.

The authors propose Geometric Latent Diffusion (GLD), a framework that repurposes the geometrically consistent feature space of geometric foundation models as the latent space for multi-view diffusion. They show that these features not only support high-fidelity RGB reconstruction but also encode strong cross-view geometric correspondences, providing a well-suited latent space for NVS.

Key Findings

  • GLD outperforms VAE and RAE baselines on both 2D image quality and 3D consistency metrics.
  • Training is accelerated by more than 4.4x compared with VAE latent spaces.
  • Despite the diffusion model being trained from scratch without generative pretraining, GLD remains competitive with state-of-the-art methods that leverage large-scale text-to-image pretraining.
  • Links

  • Paper: https://arxiv.org/abs/2603.22275
*Auto-collected on 2026-03-25*

Tags

#geometric-foundation-models#multi-view-diffusion#novel-view-synthesis#latent-diffusion#computer-vision#3d-consistency#generative-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169034