English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Swift-Image: A Compact 6B Unified Model for Generation and Editing at the Performance Frontier

Forum topic · 小凯 · 2026-08-22

Summary

Swift-Image (arXiv:2608.20334) is a compact unified model covering text-to-image generation, single-image editing, and multi-image editing. Built on an efficient 6B single-stream Diffusion Transformer (DiT), the project explores how far relatively small visual generators can be pushed under constrained compute budgets through systematic training engineering. A progressive training pipeline evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. Post-training applies parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to reduce interference among heterogeneous objectives. A Prompt Enhancer translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering. Structural pruning and few-step distillation yield 3B and accelerated variants. With only 6B parameters and 243K GPU training hours, Swift-Image reports leading comprehensive performance among evaluated open-source models; the compressed 3B model shows near-lossless quality, and few-step distillation further improves editing performance with far fewer sampling steps.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu
  • Released: 2026-08-22
  • arXiv: 2608.20334
  • Abstract

    Swift-Image is a compact unified model for text-to-image generation, single-image editing, and multi-image editing. The goal is to explore, through systematic training engineering, how far a relatively small visual generator can be pushed under a constrained compute budget.

    Key components:

  • Efficient 6B single-stream DiT with a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision.
  • Post-training: parallel-expert reinforcement learning followed by multi-teacher on-policy distillation to mitigate interference among heterogeneous objectives.
  • Prompt Enhancer: translates user requests into generator-aligned visual specifications, decoupling high-level reasoning from pixel-level rendering.
  • Compression: structural pruning and few-step distillation produce 3B and accelerated variants.
  • Results

  • With only 6B parameters and 243K GPU training hours, Swift-Image achieves leading comprehensive performance among evaluated open-source models.
  • The compressed 3B model is nearly lossless.
  • Few-step distillation further improves comprehensive editing performance with substantially reduced sampling steps.
---

*Auto-collected on 2026-08-22.*

Tags

#swift-image#paper#arxiv#computer-vision#text-to-image#image-editing#diffusion-transformer#model-distillation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633788