Paper Overview
- Field: Computer Vision (CV)
- Authors: Inclusion AI, Tiwei Bie, Haoxing Chen
- Published: 2026-04-22
- arXiv: 2604.20796
- Fully semantic discrete tokenizer: SigLIP-VQ discretizes continuous visual inputs into tokens compatible with the language backbone.
- MoE-based dLLM backbone: performs block-level masked diffusion over both text and vision inputs in a unified sequence.
- Diffusion decoder: reconstructs visual tokens into high-fidelity images.
- Prefix-aware optimizations in the backbone improve inference beyond naive parallel decoding.
- Few-step distillation in the decoder reduces image generation cost.
- Matches specialized vision-language models on multimodal understanding benchmarks.
- Strong performance on image generation and editing; supports native interleaved generation and reasoning.
- Code and models: https://github.com/inclusionAI/LLaDA2.0-Uni
Abstract (English)
We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its architecture combines a fully semantic discrete tokenizer, a MoE-based dLLM backbone, and a diffusion decoder. By discretizing continuous visual inputs via SigLIP-VQ, the model enables block-level masked diffusion for both text and vision inputs within the backbone, while the decoder reconstructs visual tokens into high-fidelity images. Inference efficiency is enhanced beyond parallel decoding through prefix-aware optimizations in the backbone and few-step distillation in the decoder. Supported by carefully curated large-scale data and a tailored multi-stage training pipeline, LLaDA2.0-Uni matches specialized VLMs on multimodal understanding while delivering strong performance in image generation and editing. Its native support for interleaved generation and reasoning establishes a promising and scalable paradigm for next-generation unified foundation models.