Key points
This report surveys the leading text-to-image (T2I) models of 2024–2026, covering architecture, performance, licensing, and ecosystems.
Background
- T2I has evolved from GANs to diffusion models (DALL·E 2021, Stable Diffusion 2022) into a mature content-production technology by 2024–2026.
- The competitive landscape: open-source models emphasize flexibility, controllability, and local deployment; closed-source models lead in stability and commercial readiness.
- FLUX (Black Forest Labs): 12B-parameter suite ([dev] non-commercial, [schnell] Apache 2.0, [pro] API). Outperformed Midjourney v6, DALL·E 3, and SD3 on Elo benchmarks at release; supports up to 10 reference images for identity/style consistency.
- Stable Diffusion 3.5 (Stability AI): Maturest community ecosystem (LoRA, ControlNet, Civitai); improved text rendering; Community License free for individuals and companies under $1M revenue.
- AuraFlow (Fal.ai): Open-source challenger claiming to beat Midjourney v6 and DALL·E 3 HD on several metrics at launch.
- Ideogram 2.0: Specialized in text rendering (English OCR accuracy 0.97), multilingual typography; Apache 2.0; nf4 quantized version runs on 6–8GB VRAM.
- Wan 2.2 (Alibaba): Open video generation (T2V-A14B, TI2V-5B); 720P/24fps on consumer GPUs; MoE architecture; Apache 2.0 but 14B variant needs 40GB+ VRAM.
- Qwen-Image 2.0 (Alibaba): 7B unified generation/editing model, native 2K output, near-commercial multilingual (Chinese/English) text rendering; Lightning distillation cuts 40 steps to 4 (~10x speedup); Apache 2.0.
- Z-Image Turbo (Alibaba): 6B fast-generation model using Decoupled-DMD distillation (8 steps, ~1s on H100, 5–10s on 16GB consumer GPUs); Apache 2.0.
- ERNIE-Image (Baidu): 8B single-stream DiT, open-sourced April 2026; SuperCLUE Chinese T2I score 76.37 (1st in China, 4th globally); LongText-Bench 0.9733 (1st among open models); runs on 24GB GPUs; Apache 2.0 with ComfyUI and Unsloth LoRA support.
- Others: GLM-Image (Zhipu AI, AR+diffusion hybrid, strong in dense Chinese text), Boogu-Image series.
- DALL·E 3 (OpenAI): Integrated in ChatGPT with GPT-4 prompt expansion; strong text embedding and prompt understanding; conservative artistic style.
- Midjourney v6: Best-in-class photorealism and artistic style (anatomy, skin, hands); slower generation; subscription/Discord only.
- Adobe Firefly: Deep Creative Cloud integration, vector graphics generation, copyright-safe commercial training data.
- Nano Banana Pro (Google): SuperCLUE global #1 at 83.73; exceptional realism and multilingual text rendering.
- GPT-Image 2.0 (OpenAI, codename "Spud"): Autoregressive multimodal architecture replacing diffusion; "thinking mode" 8-step workflow (task decomposition, retrieval, self-check); >99% character accuracy; 2K output; Image Arena top score 1512 (+242 over Nano Banana Pro).
- Imagen 4 (Google): First-class text rendering, strong spatial/contextual understanding, 2K output, Google Workspace integration; slower generation.
- Control/customization: Open models allow LoRA fine-tuning, ControlNet, dataset customization; closed models only prompt engineering.
- Data sovereignty: Open models support private on-premise deployment; closed models send data to cloud.
- Cost: Open models require significant hardware (e.g., FLUX.2 dev FP8 ~32GB VRAM; production on A100/H100); closed APIs are pay-per-image with no hardware overhead.
- Quality/stability: Closed models generally lead in complex prompt adherence, text rendering, and batch consistency; open models can match them after fine-tuning.
- Licensing/risk: Some open licenses restrict commercial use (SD Community License >$1M revenue); closed providers assume copyright risk for a fee.