Summary
A forum post introduces the paper "Modality Forcing for Scalable Spatial Generation" (arXiv:2506.10667) by Bardienus Pieter Duisterhof, Deva Ramanan, and Jeffrey Ichnowski. Text-to-image (T2I) models contain rich spatial priors such as perspective and relative scale. The paper proposes Modality Forcing, a simple and scalable post-training recipe for joint image-depth generation using a single Diffusion Transformer (DiT) trained on sparse depth data. By assigning separate noise levels to each modality, the method supports conditional and joint generation of images and depth in any permutation. Per-modality decoders enable training on sparse real-world depth and strong, generalizable depth prediction. The authors train T2I models from 370M to 3.3B parameters, showing larger models trained on more image data yield more accurate depth, demonstrating that Modality Forcing inherits T2I scalability. The strongest model competes with state-of-the-art monocular depth estimators and reduces AbsRel by 57% relative to existing joint image-depth generative models, suggesting image generation is a scalable pretraining objective for spatial awareness.
Paper Overview
Field: Computer Vision
Authors: Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski
Published: 2025-06-13
arXiv: 2506.10667
Summary
Text-to-image (T2I) models contain rich spatial priors. Synthesizing photorealistic, cluttered scenes requires an understanding of geometry, including perspective and relative scale. Prior works adapt T2I models to leverage this prior for depth prediction, but they require dense depth data and involve complex recipes.
The authors propose Modality Forcing, a simple, scalable post-training recipe for joint image-depth generation using a single DiT trained on sparse depth data. Modality Forcing enables conditional and joint generation of image and depth in any permutation by assigning separate noise levels per modality. Per-modality decoders allow training on sparse, real-world depth and achieve strong, generalizable depth prediction.
Key Findings
- Modality Forcing inherits the scalability of T2I pretraining: training a family of T2I models from scratch (370M to 3.3B parameters) shows that larger models trained on more image data produce more accurate depth.
- The strongest model competes with state-of-the-art monocular depth estimators.
- Relative to existing joint image-depth generative models, it reduces AbsRel by 57%.
- These results provide strong evidence that image generation is a scalable pretraining objective for spatial awareness.
---
*Auto-collected on 2026-06-14*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981273