Summary
AVGen-Bench (arXiv:2504.07073) is a task-driven benchmark for evaluating Text-to-Audio-Video (T2AV) generation, a rapidly emerging interface for media creation whose assessment has remained fragmented. Created by Ziwei Zhou, Zeyuan Lai, and Rui Wang, the benchmark features high-quality prompts across 11 real-world categories. It introduces a multi-granular evaluation framework combining lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling assessment that spans perceptual quality down to fine-grained semantic controllability — addressing the limitations of prior benchmarks that evaluate audio and video in isolation or rely on coarse embedding similarity. The authors' evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability in current T2AV models, with persistent failures in text rendering, speech coherence, and physical reasoning, as well as a universal breakdown in musical pitch control. Published April 2025, this work provides a more rigorous and comprehensive methodology for measuring joint audio-video generation correctness.
Paper Overview
- Field: AI
- Authors: Ziwei Zhou, Zeyuan Lai, Rui Wang
- Published: 2025-04-10
- arXiv: 2504.07073
Abstract
Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture the fine-grained joint correctness required by realistic prompts.
The authors introduce AVGen-Bench, a task-driven benchmark for T2AV generation featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, they propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability.
Key Findings
The evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including:
- Persistent failures in text rendering
- Breakdowns in speech coherence
- Weaknesses in physical reasoning
- A universal collapse in musical pitch control
Links
- arXiv paper: https://arxiv.org/abs/2504.07073
---
*Auto-collected on 2025-04-11*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169739