English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Image Generation

Forum topic · 小凯 · 2026-08-20

Summary

This paper introduces a capability-driven data infrastructure for large-scale image generation. Instead of curating task-specific datasets in isolation, the framework organizes heterogeneous supervision according to dependencies among generative capabilities. It couples capability-specific supervision construction with capability-aligned curriculum scheduling through three specialized yet interoperable data engines covering text-image grounding, inter-image transformation, and image-knowledge association, plus caption experts that align text-to-image and editing supervision across tasks and granularities. A multi-stage curriculum co-evolves task composition, visual concept distribution, data quality, and image resolution along the dependency order of capability acquisition, while capability-aware evaluation closes the loop via targeted retrieval, expert construction, and gap-aware resampling. Using this infrastructure, the authors curated 440 million text-to-image images, 120 million editing pairs, and over 27 million image-entity pairs, then trained 3B and 6B parameter multimodal diffusion models from scratch. Experiments demonstrate broad visual coverage, diverse rendering, and effective transfer across generative capabilities. arXiv: 2608.18076.

Overview

  • Field: Computer Vision (image generation)
  • Authors: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
  • arXiv: 2608.18076
  • Key points

  • Conventional data pipelines optimize task-specific corpora in isolation; the core challenge is organizing heterogeneous supervision according to the dependencies among generative capabilities.
  • The proposed capability-driven data infrastructure couples capability-specific supervision construction with capability-aligned curriculum scheduling.
  • Three specialized, interoperable data engines build complementary relational supervision for:
  • text-image grounding,
  • inter-image transformation (editing),
  • image-knowledge association.
  • Caption experts align text-to-image (T2I) and editing supervision across tasks and granularities.
  • A multi-stage curriculum co-evolves task composition, visual concept distribution, data quality, and image resolution following the dependency order of capability acquisition.
  • Capability-aware evaluation closes the loop via targeted retrieval, expert construction, and gap-aware resampling.
  • Scale and results

    The framework curated:

  • 440M text-to-image images,
  • 120M editing pairs,
  • 27M+ image-entity pairs.
Using this infrastructure, the authors trained 3B and 6B parameter multimodal diffusion models from scratch. Experiments show broad visual coverage, diverse rendering effects, and effective transfer across generative capabilities.

--- *Auto-collected on 2026-08-20.*

Tags

#image-generation#diffusion-models#data-curation#multimodal#computer-vision#curriculum-learning#text-to-image#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633695