Overview
- Field: AI / Speech Recognition
- Authors: Geeyang Tay, Wentao Ma, Jaewon Lee, Yuzhi Tang, Daniel Lee, Weisu Yin, Dongming Shen, Silin Meng, Yi Zhu, Mu Li, Alex Smola
- Published: 2026-03-26
- arXiv: 2603.25727
- Evaluations of seven widely used ASR systems reveal severe and uneven performance degradation.
- Model robustness does not transfer across languages or conditions.
- Critically, models often hallucinate plausible content that was never spoken when given partial or degraded inputs — a concrete safety risk for downstream agent behavior.
Summary
Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without diagnostic tools that isolate specific failure factors, practitioners cannot anticipate which conditions, in which languages, will cause what degree of degradation.
The authors introduce WildASR, a multilingual (four-language) diagnostic benchmark sourced entirely from real human speech. It factorizes ASR robustness along three axes:
1. Environmental degradation 2. Demographic shift 3. Linguistic diversity
Key Findings
Implications
The results show that targeted, factor-isolated evaluation is essential for understanding and improving ASR reliability in production voice-agent systems. Beyond the benchmark itself, the authors release three analysis tools to help practitioners guide deployment decisions.
---
*Auto-collected on 2026-03-28*