Summary
Swift-Image is a compact unified visual generation model covering text-to-image generation, single-image editing, and multi-image editing, introduced in arXiv paper 2608.20334 by Taihang Hu, Zhao Wang, Zuan Gao, and Tao Liu. The work explores how far a relatively small visual generator can be pushed under a constrained compute budget through systematic training engineering. The architecture uses an efficient 6B single-stream Diffusion Transformer (DiT) trained with a progressive pipeline that moves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. Post-training combines parallel-expert reinforcement learning with multi-teacher on-policy distillation to reduce interference across heterogeneous objectives, while a Prompt Enhancer translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering. Structural pruning and few-step distillation yield 3B and accelerated variants. With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading overall performance among evaluated open-source models; the compressed 3B model retains nearly full performance, and few-step distillation further improves editing quality with far fewer sampling steps.
Overview
Field: Computer Vision (CV)
Authors: Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu
Published: 2026-08-22
arXiv: 2608.20334
Abstract
Swift-Image is a compact unified model for text-to-image generation, single-image editing, and multi-image editing. The goal is to explore, through systematic training engineering under a restricted compute budget, how far a relatively small visual generator can be pushed.
Key Techniques
- Architecture: An efficient 6B single-stream DiT (Diffusion Transformer).
- Progressive training pipeline: Evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision.
- Post-training: Parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to mitigate interference between heterogeneous objectives.
- Prompt Enhancer: Translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering.
- Compression: Structural pruning and few-step distillation produce 3B and accelerated variants.
Results
- With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading overall performance among evaluated open-source models.
- The compressed 3B model retains performance with almost no loss.
- Few-step distillation further improves overall editing performance while substantially reducing sampling steps.
---
*Auto-collected on 2026-08-22. Source: arXiv:2608.20334*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633809