Summary
This paper introduces a capability-driven data infrastructure for large-scale image generation that moves beyond optimizing task-specific datasets in isolation. The framework couples capability-specific supervision construction with capability-aligned curriculum scheduling. Three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align text-to-image (T2I) and editing supervision across tasks and granularities. A multi-stage curriculum co-evolves task composition, visual concept distribution, data quality, and image resolution along the dependency order of capability acquisition, and capability-aware evaluation closes the loop through targeted retrieval, expert construction, and gap-aware resampling. Using this infrastructure, the authors curated 440 million T2I images, 120 million editing pairs, and over 27 million image-entity pairs, then trained 3B and 6B parameter multimodal diffusion models from scratch. Experiments demonstrate broad visual coverage, diverse rendering effects, and effective transfer across generation capabilities. Paper arXiv:2608.18076.
Paper Overview
- Field: Computer Vision
- arXiv: 2608.18076
- Authors: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
Key Idea
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. The central challenge addressed here is not only how to curate each task-specific corpus, but how to organize heterogeneous supervision according to the dependencies among generative capabilities.
Approach
The authors present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling:
- Three specialized, interoperable data engines build complementary relational supervision for:
- Text-image grounding
- Inter-image transformation
- Image-knowledge association
- Caption experts align text-to-image (T2I) and editing supervision across tasks and granularities.
- Multi-stage curriculum co-evolves task composition, visual concept distribution, data quality, and image resolution along the dependency order of capability acquisition.
- Capability-aware evaluation closes the loop through targeted retrieval, expert construction, and gap-aware resampling.
Scale and Results
The framework curated:
- 440 million text-to-image images
- 120 million editing pairs
- Over 27 million image-entity pairs
Using this infrastructure, the team trained
3B and 6B parameter multimodal diffusion models from scratch. Experiments show broad visual coverage, diverse rendering effects, and effective transfer across generation capabilities.
---
*Source: forum post, arXiv 2608.18076.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633673