English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

FlatSounds: Benchmarking Single-Factor Physical Video-to-Audio Generation

Forum topic · 小凯 · 2026-06-01

Summary

This paper introduces FlatSounds, a benchmark for auditing generative video-to-audio (V2A) models on physical reasoning. While current V2A models produce highly plausible soundtracks, existing evaluations emphasize perceptual realism and overlook physical correctness under controlled interventions. FlatSounds provides: (1) controlled counterfactual pairs where a single physical factor varies, and (2) single-video pattern tests probing internal consistency and directional trends, checking whether generated audio correctly reflects specific physical properties and timing. Evaluations of state-of-the-art models reveal a consistent trade-off: models rely more on text captions than visual streams to infer physics and semantics. Captions generally improve physical and semantic accuracy but paradoxically degrade temporal alignment. The authors highlight the need to move beyond audio quality toward learning physical processes directly from pixels. They also show that physics-based metrics correlate highly with human preference tests on their own data. Posted on zhichai.net with arXiv link 2605.30339.

FlatSounds: Benchmarking Single-Factor Physical Video-to-Audio Generation

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Tingle Li, Siddharth Gururani, Kevin J. Shih, Gantavya Bhatt, Sang-gil Lee, Zhifeng Kong, Arushi Goel, Gopala Anumanchipalli, Ming-Yu Liu
  • Published: 2026-05-28
  • arXiv: 2605.30339
  • Summary

    Generative video-to-audio (V2A) models produce highly plausible soundtracks, but whether they capture the underlying physical processes remains unclear. Existing evaluations emphasize perceptual realism while ignoring physical correctness under controlled interventions.

    This paper introduces the FlatSounds benchmark to audit the physical reasoning of V2A models through:

    1. Controlled counterfactual pairs with variation in a single physical factor; 2. Single-video pattern tests probing internal consistency and directional trends.

    These settings test whether the generated audio correctly reflects specific physical properties and timing.

    Key Findings

  • Evaluation of state-of-the-art models reveals a consistent trade-off: models rely more on text captions than visual streams to infer physics and semantics.
  • Captions generally improve physical and semantic accuracy, but paradoxically degrade temporal alignment.
  • The results highlight the need to move from audio quality toward learning physical processes directly from pixels.
  • Physics-based metrics correlate highly with the authors' own human preference tests on their data.
---

*Auto-collected on 2026-06-01.*

Tags

#video-to-audio#benchmark#generative-models#physics#computer-vision#audio-generation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980678