English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Probing Automated TTS Evaluators on Linguistically Grounded Dimensions

Forum topic · 小凯 · 2026-08-11

Summary

Automated text-to-speech (TTS) evaluation methods, including Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges, are expected to mirror human perception, but how well they capture the distinct aspects of speech listeners actually perceive remains unclear. This paper (arXiv:2508.03806) deconstructs 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions and builds the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Benchmarking four MOS predictors and four Audio-LLM judges shows that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges exhibit selective, prompt-dependent detection that fails to generalize across all dimensions. Neither approach reliably captures broad linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

Paper Overview

Research Area: NLP Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah arXiv: 2508.03806

Abstract

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. The authors deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.

Key Findings

  • MOS predictors collapse onto acoustic signal quality — they do not differentiate finer linguistic dimensions of perceived speech quality.
  • Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions.
  • Neither family of evaluators reliably captures broad linguistically structured speech errors.
  • The dataset, annotation framework, and evaluation code are publicly released to enable more targeted and interpretable TTS evaluation.

Significance

This work provides the first dimension-level meta-evaluation benchmark for TTS, shifting evaluation beyond a single 'naturalness' score toward fine-grained, linguistically meaningful assessment of synthetic speech.

---

Source: arXiv:2508.03806

Tags

#tts#speech-evaluation#nlp#audio-llm#mos-prediction#benchmark#naturalness

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633356