Summary
GPIC (Giant Permissive Image Corpus) is a large-scale dataset for visual generation research introduced by researchers including Keshigeyan Chandrasegaran and Li Fei-Fei. The corpus contains approximately 28 trillion pixels of diverse internet imagery, annotated with state-of-the-art vision-language models. It includes 100 million training samples, 200,000 validation samples, and 1 million test samples, all licensed for both research and commercial use. The dataset has undergone safety filtering and deduplication, and is centrally hosted on Hugging Face. Alongside the corpus, the authors provide a generative modeling benchmark protocol and a reference baseline based on pixel-space flow matching. GPIC addresses the need for large, accessible, and stable datasets to support scalable visual generation modeling research. The dataset, benchmark, and models are available on Hugging Face.
Paper Overview
Research Area: Computer Vision (CV)
Authors: Keshigeyan Chandrasegaran, Kyle Sargent, Suchir Agarwal, Michael Jang, Michael Poli, Juan Carlos Niebles, Justin Johnson, Jiajun Wu, Li Fei-Fei
arXiv: 2605.30341
Abstract
Scalable approaches to visual generation modeling require large, accessible, and stable datasets. The authors introduce GPIC, a Giant Permissive Image Corpus comprising approximately 28 trillion pixels.
Key properties of the corpus:
- Diverse internet images, annotated with state-of-the-art vision-language models
- 100 million training samples, 200,000 validation samples, and 1 million test samples
- All images are licensed for both research and commercial use
- Safety-filtered and deduplicated
- Centrally hosted on Hugging Face
Alongside the dataset, the authors provide:
- A generative modeling benchmark protocol for GPIC
- A reference baseline based on pixel-space flow matching
The dataset, benchmark, and models are all available on Hugging Face.
---
*Automatically collected on 2026-06-01.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177980677