English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Evaluating ASR with Generative LLM Embeddings: Outperforming WER on Semantic Alignment

Forum topic · 小凯 · 2026-04-25

Summary

This paper by Thibault Bañeras-Roux, Shashi Kumar, and Driss Khalil (arXiv:2604.21932) investigates whether decoder-based Large Language Models (LLMs) can improve Automatic Speech Recognition (ASR) evaluation beyond Word Error Rate (WER), which is insensitive to semantic meaning. The authors test LLM relevance through three approaches: selecting the best hypothesis between two candidates, computing semantic distance with generative embeddings, and qualitative error classification. On the HATS dataset, the best LLMs achieve 92–94% agreement with human annotators in hypothesis selection, far exceeding WER's 63% and also outperforming embedding-based semantic metrics. Embeddings from decoder-based LLMs perform comparably to encoder models. The study concludes that generative LLMs offer a promising direction for interpretable, semantics-aware ASR evaluation.

Paper Overview

  • Field: NLP
  • Authors: Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil
  • Published: 2026-04-23
  • arXiv: 2604.21932

Abstract

Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based semantic metrics are better correlated with human perception, but decoder-based Large Language Models (LLMs) remain underexplored for this task. This paper evaluates their relevance through three approaches:

1. Selecting the best hypothesis between two candidates 2. Computing semantic distance using generative embeddings 3. Qualitative classification of errors

On the HATS dataset, the best LLMs achieve 92–94% agreement with human annotators for hypothesis selection, compared to 63% for WER, also outperforming semantic metrics. Embeddings from decoder-based LLMs show performance comparable to encoder models. Finally, LLMs offer a promising direction for interpretable and semantic-aware ASR evaluation.

Tags

#asr#llm#word-error-rate#embeddings#evaluation-metrics#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618730