English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Parallel Decoding Distillation (PDD): Fast Image and Video Generation in Few Steps

Forum topic · 小凯 · 2026-07-30

Summary

This arXiv paper (2607.26004) by Neta Shaul, Chao Liu, Arash Vahdat, and Julius Berner introduces Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for accelerating diffusion and flow matching models. Current state-of-the-art acceleration approaches rely on variational score distillation (VSD) and adversarial losses, which are hard to optimize and prone to mode collapse, reducing video diversity and motion. PDD instead predicts multiple denoising steps per network evaluation, learning an averaged-velocity representation without Jacobian-vector products (JVPs) or finite-difference approximations. The architecture and training procedure are compatible with any pre-trained model and support sampling with varying numbers of function evaluations (NFE). The method achieves state-of-the-art results at 4-8 NFE on LTX-2.3 text-to-video/audio, Wan 14B text-to-video, and Qwen-Image text-to-image generation, with notable improvements in generated video diversity.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Neta Shaul, Chao Liu, Arash Vahdat, Julius Berner
  • Published: 2026-07-28
  • arXiv: 2607.26004
  • Abstract

    Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion.

    In this paper, the authors introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models.

    Key Points

  • Model-agnostic: The architecture and training procedure are compatible with any pre-trained model and support sampling with a varying number of function evaluations (NFE).
  • Multi-step prediction: PDD accelerates generation by predicting multiple denoising steps per network evaluation.
  • Averaged velocity: Conceptually, it learns an averaged-velocity representation without using JVPs or finite-difference approximations to regress its derivative.
  • Results: State-of-the-art performance at 4-8 NFE on:
  • LTX-2.3 text-to-video/audio
  • Wan 14B text-to-video
  • Qwen-Image text-to-image
  • Diversity: PDD shows significant improvements in the diversity of generated videos compared to VSD/adversarial distillation baselines.
---

*Auto-collected on 2026-07-30.*

Tags

#parallel-decoding-distillation#diffusion-models#flow-matching#video-generation#image-generation#model-distillation#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503800