English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Probing Automated TTS Evaluators on Linguistically Grounded Dimensions

Forum topic · 小凯 · 2026-08-11

Summary

Automated text-to-speech (TTS) evaluation methods—Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges—are expected to reflect human perception, but how well they capture the distinct aspects of speech listeners actually perceive is unclear. This paper deconstructs 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions and uses it to build the first dimension-level meta-evaluation benchmark for TTS, consisting of 860 utterances annotated by trained linguist raters. Benchmarking four MOS predictors and four Audio-LLM judges shows that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges exhibit selective, prompt-dependent detection that fails to generalize across all dimensions. Neither approach reliably captures a broad range of linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

Paper Overview

  • Research area: NLP
  • Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah
  • Published: 2026-08-11
  • arXiv: 2508.03806
  • Summary

    Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.

    The authors deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters.

    Key Findings

  • MOS predictors: Collapse onto acoustic signal quality rather than capturing diverse perceptual dimensions.
  • Audio-LLM judges: Show selective, prompt-dependent detection that does not generalize across all dimensions.
  • Overall: Neither approach reliably captures a broad range of linguistically structured speech errors.
The dataset, annotation framework, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

--- *Automatically collected on 2026-08-12*

Tags

#tts#speech-evaluation#nlp#audio-llm#mos-prediction#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633334