Summary
Swift-Image (arXiv:2608.20334) is a compact unified model covering text-to-image generation, single-image editing, and multi-image editing. Built on an efficient 6B single-stream Diffusion Transformer (DiT), the project explores how far relatively small visual generators can be pushed under constrained compute budgets through systematic training engineering. A progressive training pipeline evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. Post-training applies parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to reduce interference among heterogeneous objectives. A Prompt Enhancer translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering. Structural pruning and few-step distillation yield 3B and accelerated variants. With only 6B parameters and 243K GPU training hours, Swift-Image reports leading comprehensive performance among evaluated open-source models; the compressed 3B model shows near-lossless quality, and few-step distillation further improves editing performance with far fewer sampling steps.
Paper Overview
- Field: Computer Vision (CV)
- Authors: Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu
- Released: 2026-08-22
- arXiv: 2608.20334
Abstract
Swift-Image is a compact unified model for text-to-image generation, single-image editing, and multi-image editing. The goal is to explore, through systematic training engineering, how far a relatively small visual generator can be pushed under a constrained compute budget.
Key components:
- Efficient 6B single-stream DiT with a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision.
- Post-training: parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to mitigate interference among heterogeneous objectives.
- Prompt Enhancer: translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering.
- Compression: structural pruning and few-step distillation produce 3B and accelerated variants.
Results
- With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading comprehensive performance among evaluated open-source models.
- The compressed 3B model is nearly lossless.
- Few-step distillation further improves comprehensive editing performance with substantially reduced sampling steps.
---
*Auto-collected on 2026-08-22.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633788