Paper Overview
Research Area: NLP Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah arXiv: 2508.03806
Abstract
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. The authors deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.
Key Findings
- MOS predictors collapse onto acoustic signal quality — they do not differentiate finer linguistic dimensions of perceived speech quality.
- Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions.
- Neither family of evaluators reliably captures broad linguistically structured speech errors.
- The dataset, annotation framework, and evaluation code are publicly released to enable more targeted and interpretable TTS evaluation.
Significance
This work provides the first dimension-level meta-evaluation benchmark for TTS, shifting evaluation beyond a single 'naturalness' score toward fine-grained, linguistically meaningful assessment of synthetic speech.
---
Source: arXiv:2508.03806