English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Forum topic · 小凯 · 2026-04-12

Summary

AVGen-Bench (arXiv:2504.07857) is a task-driven benchmark for evaluating Text-to-Audio-Video (T2AV) generation, introduced by Ziwei Zhou, Zeyuan Lai, and Rui Wang (posted 2025-04-10). Existing T2AV evaluation is fragmented: prior benchmarks assess audio and video in isolation or rely on coarse embedding similarity, failing to capture fine-grained joint correctness demanded by realistic prompts. AVGen-Bench provides high-quality prompts spanning 11 real-world categories and a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), measuring everything from perceptual quality to fine-grained semantic controllability. Evaluations reveal a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including persistent failures in text rendering, speech coherence, and physical reasoning, as well as a widespread inability to control music pitch. This positions AVGen-Bench as a comprehensive standard for diagnosing both quality and controllability in next-generation T2AV models.

Paper Overview

Research Area: NLP Authors: Ziwei Zhou, Zeyuan Lai, Rui Wang Published: 2025-04-10 arXiv: 2504.07857

Abstract (English)

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture the fine-grained joint correctness required by realistic prompts.

The authors introduce AVGen-Bench, a task-driven benchmark for T2AV generation featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, they propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability.

Key Findings

  • Significant gap between strong audio-visual aesthetics and weak semantic reliability in current T2AV models
  • Persistent failures in:
  • Text rendering
  • Speech coherence
  • Physical reasoning
  • Widespread inability to control music pitch

Why It Matters

By unifying perceptual and semantic evaluation at multiple granularities, AVGen-Bench offers a more faithful measure of real-world T2AV utility, helping researchers identify precise failure modes rather than relying on coarse similarity metrics.

--- *Auto-collected on 2026-04-12*

Tags

#text-to-audio-video#benchmark#multimodal#t2av-generation#mlLM#evaluation-framework#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169765