Paper Overview
Research Field: Machine Learning (ML)
Authors: Geeyang Tay, Wentao Ma, Jaewon Lee, Yuzhi Tang, Daniel Lee, et al.
Published: 2026-03-26
arXiv: 2603.25727
Abstract
Automatic Speech Recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet they still fail in real-world voice agents. This paper introduces WildASR, a fully multilingual (four-language) diagnostic benchmark sourced entirely from genuine human speech. The benchmark decomposes ASR robustness along three axes:
1. Environmental degradation — robustness to noise, channel distortion, and acoustic conditions 2. Demographic shift — robustness across speaker demographics (age, gender, accent) 3. Linguistic diversity — robustness across multiple languages
Key Findings
- Evaluation reveals severe and uneven performance degradation across the three robustness axes.
- Model robustness does not transfer across languages or across conditions — high accuracy in one setting does not imply reliable performance in another.
- The gap between curated benchmark scores and real-world voice agent reliability remains substantial, motivating robustness-focused evaluation methodologies.
Implications
The work highlights a critical disconnect between standard ASR benchmarks and deployment scenarios involving voice agents. WildASR provides a reproducible testbed for measuring practical ASR robustness and exposes failure modes that aggregate metrics obscure.
--- *Auto-collected 2026-03-29*