English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Lip Forcing: Few-Step Autoregressive Diffusion for Real-Time Lip Synchronization

Forum topic · 小凯 · 2026-06-11

Summary

Lip Forcing is presented as the first autoregressive diffusion method for video-to-video (V2V) lip synchronization. The approach distills a 14B-parameter audio-conditioned bidirectional video diffusion teacher into causal student models. At inference, students generate each chunk with only two denoising steps and without inference-time classifier-free guidance (CFG), enabling real-time lip synchronization. A lip-sync-specific teacher-trajectory analysis reveals a CFG fidelity-sync tradeoff: no-CFG predictions favor reference fidelity, while CFG-guided predictions favor synchronization within a mid-trajectory band. The 1.3B student model achieves 31 FPS, running 17.6x faster than a bidirectional model of similar size, and the 14B student is 39.8x faster than its teacher, with sub-millisecond first-frame latency. The paper (arXiv 2606.11180) was released on June 9, 2026, in the computer vision field.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee, Siyoon Jin, Heeseong Shin, Jung Yi, Yunjin Park, Chulmin Park, Seungryong Kim
  • Released: 2026-06-09
  • arXiv: 2606.11180
  • Abstract

    Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. The authors present Lip Forcing, to their knowledge the first autoregressive diffusion method for video-to-video (V2V) lip synchronization, which distills a 14B audio-conditioned bidirectional video diffusion teacher into causal students.

    At inference, the students generate each chunk in only two denoising steps without inference-time CFG, enabling real-time lip synchronization.

    A lip-sync-specific teacher-trajectory analysis reveals a CFG fidelity-sync tradeoff: no-CFG predictions favor reference fidelity, whereas CFG-guided predictions favor synchronization within a mid-trajectory band.

    Key Results

  • 1.3B student model: 31 FPS, 17.6x faster than a bidirectional model of the same scale
  • 14B student model: 39.8x faster than the teacher
  • Sub-millisecond first-frame latency

Summary (Chinese Forum Abstract)

Lip Forcing is the first autoregressive diffusion approach for V2V lip synchronization, distilling a 14B audio-conditioned bidirectional video diffusion teacher into a causal student. With only 2 denoising steps at inference and no inference-time CFG, it achieves real-time lip synchronization.

---

*Auto-collected on 2026-06-11*

Tags

#lip-synchronization#diffusion-models#autoregressive#video-generation#model-distillation#real-time#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981077