English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Probing Automated TTS Evaluators on Linguistically Grounded Dimensions

Forum topic · 小凯 · 2026-08-11

Summary

This paper investigates how well automated Text-to-Speech (TTS) evaluation methods align with the specific aspects of speech that human listeners perceive. The authors deconstruct the monolithic concept of 'naturalness' into a linguistically grounded annotation schema covering 10 distinct perceptual dimensions. Using this schema, they construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Benchmarking four MOS predictors and four Audio-LLM judges reveals two failure modes: MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that fails to generalise across all dimensions. Neither family reliably captures the broader range of linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to enable more targeted and interpretable TTS evaluation.

Beyond Naturalness: Probing Automated TTS Evaluators on Linguistically Grounded Dimensions

Field: NLP Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah Published: 2026-08-11 arXiv: 2508.03806

Summary

Automated Text-to-Speech (TTS) evaluation methods—Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges—are expected to reflect human perception. However, it remains unclear how well they capture the distinct aspects of speech that listeners actually perceive. The authors deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions. Using this schema, they construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.

Key Findings

  • MOS predictors collapse onto acoustic signal quality, failing to discriminate among the broader set of perceptually relevant dimensions.
  • Audio-LLM judges exhibit selective, prompt-dependent detection that does not generalise across all dimensions.
  • Neither evaluation paradigm reliably captures the wide range of linguistically structured speech errors.
  • Contributions

  • A 10-dimension linguistically grounded annotation schema that operationalises what listeners actually perceive in synthesised speech.
  • A dimension-level meta-evaluation benchmark (860 utterances, expert-annotated).
  • An empirical comparison of four MOS predictors and four Audio-LLM judges against this benchmark.
  • Public release of the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.
  • Original Abstract

    > Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dim…

    Links

  • Paper: https://arxiv.org/abs/2508.03806

Tags

#tts#text-to-speech#evaluation#mos-predictor#audio-llm#nlp#benchmark#linguistic-annotation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633347