English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CubiD: Cubic Discrete Diffusion for High-Dimensional Visual Generation

Forum topic · 小凯 · 2026-03-22

Summary

CubiD (Cubic Discrete Diffusion) is introduced as the first discrete generation model designed for high-dimensional representations. Instead of operating on low-dimensional token sequences, CubiD performs fine-grained masking throughout a high-dimensional discrete representation, enabling discrete diffusion to scale to demanding visual generation tasks. On the ImageNet-256 benchmark, CubiD achieves state-of-the-art discrete generation performance and demonstrates strong scaling behavior across model sizes from 900M to 3.7B parameters. Notably, the same discrete tokens can effectively serve both understanding and generation tasks, unifying the two capabilities within a single representation framework. The work was posted to arXiv as paper 2503.16927 (CV research area) by Yuqing Wang, Chuofan Ma, and Zhijie Lin. This article summarizes the paper's core contribution: bringing discrete diffusion to high-dimensional representations via cubic fine-grained masking, and its implications for unified vision models.

Paper Overview

Research Area: Computer Vision (CV) Authors: Yuqing Wang, Chuofan Ma, Zhijie Lin arXiv: 2503.16927

Abstract

We present Cubic Discrete Diffusion (CubiD), the first discrete generation model for high-dimensional representations. CubiD performs fine-grained masking throughout the high-dimensional discrete representation. On ImageNet-256, CubiD achieves state-of-the-art discrete generation with strong scaling behavior from 900M to 3.7B parameters.

Key Points

  • First of its kind: CubiD is the first discrete generation model that operates directly on high-dimensional representations, extending discrete diffusion beyond low-dimensional token sequences.
  • Fine-grained masking: The method applies masking operations at a fine granularity across the entire high-dimensional discrete representation, enabling more expressive denoising during generation.
  • State-of-the-art results: On ImageNet-256, CubiD sets a new state of the art for discrete generation models.
  • Strong scaling: The model exhibits healthy scaling behavior across parameter counts ranging from 900M to 3.7B.
  • Unified tokens: The same discrete tokens can effectively support both understanding and generation tasks, pointing toward unified visual models built on a single discrete representation.
  • Links

  • arXiv page: https://arxiv.org/abs/2503.16927

Tags

#discrete-diffusion#image-generation#cv#imagenet#arxiv#generative-models#scaling-laws

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168980