English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WavFlow: Audio Generation Directly in Raw Waveform Space

Forum topic · 小凯 · 2026-05-20

Summary

WavFlow is a framework that challenges the dominant latent-space compression paradigm in audio generation by synthesizing high-fidelity audio directly in raw waveform space. To handle high-dimensional low-energy signals, the authors reshape waveforms into 2D token grids via waveform patchify and apply amplitude lifting to align signal scales, enabling stable optimization through direct x-prediction in flow matching. They also curate 5 million high-quality video-text-audio triplets through an automated pipeline to learn fine-grained semantic alignment and temporal synchronization from scratch. On the VGGSound video-to-audio benchmark, WavFlow achieves FD_PaSST 59.98, IS_PANNs 17.40, and DeSync 0.44. On the AudioCaps text-to-audio benchmark, it reaches FD_PANNs 10.63 and IS_PANNs 12.62, matching or surpassing latent-space baselines. The results demonstrate that intermediate compression is not a prerequisite for high-quality synthesis, offering a simpler and more scalable alternative for multimodal audio generation.

Paper Overview

  • Research field: Computer Vision (CV)
  • Authors: Feiyan Zhou, Luyuan Wang, Shoufa Chen
  • Release date: 2026-05-19
  • arXiv: 2505.14308
  • Summary

    Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, the authors challenge this paradigm with WavFlow, a framework that generates high-fidelity audio directly in raw waveform space without intermediate representations.

    Key Contributions

  • Waveform patchify: Reshapes audio into 2D token grids to overcome the inherent difficulties of modeling high-dimensional, low-energy signals.
  • Amplitude lifting: Aligns signal scales, enabling stable optimization via direct x-prediction in flow matching.
  • Large-scale data pipeline: An automated pipeline curates 5 million high-quality video-text-audio triplets, allowing the model to learn fine-grained acoustic patterns from scratch.

Experimental Results

| Benchmark | Metric | Score | |---|---|---| | VGGSound (V2A) | FD_PaSST | 59.98 | | VGGSound (V2A) | IS_PANNs | 17.40 | | VGGSound (V2A) | DeSync | 0.44 | | AudioCaps (T2A) | FD_PANNs | 10.63 | | AudioCaps (T2A) | IS_PANNs | 12.62 |

WavFlow achieves competitive or superior performance compared with existing latent-space methods on both video-to-audio (VGGSound) and text-to-audio (AudioCaps) benchmarks. The work demonstrates that intermediate compression is not a prerequisite for high-quality synthesis, providing a simpler and more scalable alternative for multimodal audio generation.

Tags

#wavflow#audio-generation#waveform-space#flow-matching#video-to-audio#text-to-audio#vggsound#audiocaps

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620482