English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Naturalness: Probing TTS Evaluators on Linguistically Grounded Dimensions

Forum topic · 小凯 · 2026-08-11

Summary

This paper investigates how well automated Text-to-Speech (TTS) evaluation methods capture the distinct perceptual aspects of synthetic speech. The authors deconstruct the vague concept of 'naturalness' into a linguistically grounded annotation schema covering 10 perceptual dimensions and construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Benchmarking four Mean Opinion Score (MOS) predictors and four Audio-LLM judges shows that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges exhibit selective, prompt-dependent detection that does not generalize across dimensions. Neither approach reliably captures a wide range of linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

Paper Overview

  • Research area: NLP
  • Authors: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah
  • Release date: 2026-08-11
  • arXiv: 2508.03806

Chinese Summary

Automated Text-to-Speech (TTS) evaluation methods—Mean Opinion Score (MOS) predictors and Audio Large Language Model (Audio-LLM) judges—are expected to reflect human perception, yet it remains unclear how well they capture the distinct aspects of speech that listeners actually perceive. The authors deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Benchmarking four MOS predictors and four Audio-LLM judges reveals that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalize across all dimensions. Neither class of methods reliably captures a broad range of linguistically structured speech errors. The authors publicly release the dataset, annotation framework, and evaluation code to support more targeted and interpretable TTS evaluation.

Original Abstract

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct 'naturalness' into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dim...

--- *Auto-collected 2026-08-12*

#paper #arXiv #NLP

Tags

#text-to-speech#tts-evaluation#naturalness#mos-prediction#audio-llm#linguistic-annotation#meta-evaluation#benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633356