English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Swift-Image: A Compact 6B Unified Model for Image Generation and Editing

Forum topic · 小凯 · 2026-08-22

Summary

Swift-Image is a compact unified visual generation model covering text-to-image generation, single-image editing, and multi-image editing, introduced in arXiv paper 2608.20334 by Taihang Hu, Zhao Wang, Zuan Gao, and Tao Liu. The work explores how far a relatively small visual generator can be pushed under a constrained compute budget through systematic training engineering. The architecture uses an efficient 6B single-stream Diffusion Transformer (DiT) trained with a progressive pipeline that moves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. Post-training combines parallel-expert reinforcement learning with multi-teacher on-policy distillation to reduce interference across heterogeneous objectives, while a Prompt Enhancer translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering. Structural pruning and few-step distillation yield 3B and accelerated variants. With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading overall performance among evaluated open-source models; the compressed 3B model retains nearly full performance, and few-step distillation further improves editing quality with far fewer sampling steps.

Overview

Field: Computer Vision (CV) Authors: Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu Published: 2026-08-22 arXiv: 2608.20334

Abstract

Swift-Image is a compact unified model for text-to-image generation, single-image editing, and multi-image editing. The goal is to explore, through systematic training engineering under a restricted compute budget, how far a relatively small visual generator can be pushed.

Key Techniques

  • Architecture: An efficient 6B single-stream DiT (Diffusion Transformer).
  • Progressive training pipeline: Evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision.
  • Post-training: Parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to mitigate interference between heterogeneous objectives.
  • Prompt Enhancer: Translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering.
  • Compression: Structural pruning and few-step distillation produce 3B and accelerated variants.
  • Results

  • With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading overall performance among evaluated open-source models.
  • The compressed 3B model retains performance with almost no loss.
  • Few-step distillation further improves overall editing performance while substantially reducing sampling steps.
---

*Auto-collected on 2026-08-22. Source: arXiv:2608.20334*

Tags

#swift-image#text-to-image#image-editing#diffusion-transformer#model-compression#distillation#computer-vision#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633809