Beyond Naturalness: Probing Automated TTS Evaluators on Linguistically Grounded Dimensions
Field: NLP Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah Published: 2026-08-11 arXiv: 2508.03806
Summary
Automated Text-to-Speech (TTS) evaluation methods—Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges—are expected to reflect human perception. However, it remains unclear how well they capture the distinct aspects of speech that listeners actually perceive. The authors deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions. Using this schema, they construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.
Key Findings
- MOS predictors collapse onto acoustic signal quality, failing to discriminate among the broader set of perceptually relevant dimensions.
- Audio-LLM judges exhibit selective, prompt-dependent detection that does not generalise across all dimensions.
- Neither evaluation paradigm reliably captures the wide range of linguistically structured speech errors.
- A 10-dimension linguistically grounded annotation schema that operationalises what listeners actually perceive in synthesised speech.
- A dimension-level meta-evaluation benchmark (860 utterances, expert-annotated).
- An empirical comparison of four MOS predictors and four Audio-LLM judges against this benchmark.
- Public release of the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.
- Paper: https://arxiv.org/abs/2508.03806
Contributions
Original Abstract
> Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dim…