Paper Overview
Research Area: Computer Vision (CV) Authors: Yuqing Wang, Chuofan Ma, Zhijie Lin arXiv: 2503.16927
Abstract
We present Cubic Discrete Diffusion (CubiD), the first discrete generation model for high-dimensional representations. CubiD performs fine-grained masking throughout the high-dimensional discrete representation. On ImageNet-256, CubiD achieves state-of-the-art discrete generation with strong scaling behavior from 900M to 3.7B parameters.
Key Points
- First of its kind: CubiD is the first discrete generation model that operates directly on high-dimensional representations, extending discrete diffusion beyond low-dimensional token sequences.
- Fine-grained masking: The method applies masking operations at a fine granularity across the entire high-dimensional discrete representation, enabling more expressive denoising during generation.
- State-of-the-art results: On ImageNet-256, CubiD sets a new state of the art for discrete generation models.
- Strong scaling: The model exhibits healthy scaling behavior across parameter counts ranging from 900M to 3.7B.
- Unified tokens: The same discrete tokens can effectively support both understanding and generation tasks, pointing toward unified visual models built on a single discrete representation.
- arXiv page: https://arxiv.org/abs/2503.16927