English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Probing Automated TTS Evaluators with a Dimension-Level Benchmark

Forum topic · 小凯 · 2026-08-12

Summary

A new paper (arXiv:2508.05162) questions how well automated Text-to-Speech (TTS) evaluation methods capture what human listeners actually perceive. The authors deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions and build the first dimension-level meta-evaluation benchmark for TTS, consisting of 860 utterances annotated by trained linguist raters. Benchmarking four MOS (Mean Opinion Score) predictors and four Audio-LLM judges reveals key shortcomings: MOS predictors collapse onto acoustic signal quality alone, while Audio-LLM judges show selective, prompt-dependent detection that fails to generalize across all dimensions. Neither approach reliably captures linguistically structured speech errors. The team publicly releases the dataset, annotation framework, and evaluation code to enable more targeted and interpretable TTS evaluation.

Paper Overview

  • Field: NLP
  • Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah
  • Published: 2026-08-12
  • arXiv: 2508.05162
  • Summary

    Automated Text-to-Speech (TTS) evaluation methods — including Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges — are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.

    This work deconstructs "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions. Using this schema, the authors construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.

    Key Findings

  • Benchmarking covers four MOS predictors and four Audio-LLM judges.
  • MOS predictors collapse onto acoustic signal quality, ignoring deeper linguistic dimensions.
  • Audio-LLM judges show selective, prompt-dependent detection that does not generalize across all dimensions.
  • Neither evaluation approach reliably captures linguistically structured speech errors.

Resources

The dataset, annotation framework, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

--- *Auto-collected on 2026-08-12*

Tags

#tts#nlp#speech-evaluation#audio-llm#mos-predictors#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633370