English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Forum topic · 小凯 · 2026-04-11

Summary

AVGen-Bench (arXiv:2504.07073) is a task-driven benchmark for evaluating Text-to-Audio-Video (T2AV) generation, a rapidly emerging interface for media creation whose assessment has remained fragmented. Created by Ziwei Zhou, Zeyuan Lai, and Rui Wang, the benchmark features high-quality prompts across 11 real-world categories. It introduces a multi-granular evaluation framework combining lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling assessment that spans perceptual quality down to fine-grained semantic controllability — addressing the limitations of prior benchmarks that evaluate audio and video in isolation or rely on coarse embedding similarity. The authors' evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability in current T2AV models, with persistent failures in text rendering, speech coherence, and physical reasoning, as well as a universal breakdown in musical pitch control. Published April 2025, this work provides a more rigorous and comprehensive methodology for measuring joint audio-video generation correctness.

Paper Overview

  • Field: AI
  • Authors: Ziwei Zhou, Zeyuan Lai, Rui Wang
  • Published: 2025-04-10
  • arXiv: 2504.07073
  • Abstract

    Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture the fine-grained joint correctness required by realistic prompts.

    The authors introduce AVGen-Bench, a task-driven benchmark for T2AV generation featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, they propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability.

    Key Findings

    The evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including:

  • Persistent failures in text rendering
  • Breakdowns in speech coherence
  • Weaknesses in physical reasoning
  • A universal collapse in musical pitch control
  • Links

  • arXiv paper: https://arxiv.org/abs/2504.07073
---

*Auto-collected on 2025-04-11*

Tags

#text-to-audio-video#benchmark#multimodal-ai#generative-models#evaluation-framework#mlLM#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169739