English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Evaluating Automated TTS Metrics Across 10 Perceptual Dimensions

Forum topic · 小凯 · 2026-08-12

Summary

This paper examines how well automated Text-to-Speech (TTS) evaluation methods capture the multidimensional aspects of speech that human listeners actually perceive. The authors deconstruct the vague concept of 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, then construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Benchmarking four Mean Opinion Score (MOS) predictors and four Audio Large Language Model (Audio-LLM) judges reveals that MOS predictors collapse onto acoustic signal quality alone, while Audio-LLM judges exhibit selective, prompt-dependent detection that fails to generalize across all dimensions. Neither approach reliably captures linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

Paper Overview

  • Field: Natural Language Processing (NLP)
  • Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah
  • Published: 2026-08-12
  • arXiv: 2508.05162
  • Original Abstract

    Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither approach reliably captures linguistically structured speech errors. We publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

    Key Points

  • Problem: Existing TTS evaluation methods conflate multiple perceptual qualities under the umbrella term 'naturalness,' obscuring whether they truly reflect human listening judgments.
  • Framework: The authors introduce a linguistically grounded annotation schema covering 10 distinct perceptual dimensions, replacing the monolithic notion of naturalness with fine-grained criteria.
  • Benchmark: They release the first dimension-level meta-evaluation benchmark for TTS, containing 860 utterances annotated by trained linguist raters.
  • MOS Predictors: Tested across four MOS predictors, results show they collapse onto acoustic signal quality, failing to differentiate among linguistic dimensions.
  • Audio-LLM Judges: Four Audio-LLM judges demonstrate selective, prompt-dependent detection capabilities that do not generalize across all 10 dimensions.
  • Implication: Neither MOS predictors nor Audio-LLM judges reliably capture linguistically structured speech errors, motivating more targeted and interpretable TTS evaluation pipelines.
  • Release: Dataset, annotation framework, and evaluation code are publicly available to the research community.

Tags

#tts#speech-synthesis#evaluation#mos-prediction#audio-llm#nlp#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633370