Paper Overview
Field: NLP Authors: Ruohan Liu, Shukang Yin, Tao Wang Released: 2026-04-22 arXiv: 2604.20842
Introduction
Paralinguistic cues—the non-verbal aspects of speech such as tone, emotion, and prosody—are essential for natural human-computer interaction. However, evaluating these capabilities in Large Audio-Language Models (LALMs) has been limited by coarse feature coverage and the inherent subjectivity of assessment methods.
SpeechParaling-Bench
To address these challenges, the authors introduce SpeechParaling-Bench, a comprehensive benchmark for paralinguistic-aware speech generation with the following features:
- Expanded feature coverage: Grows from fewer than 50 to over 100 fine-grained paralinguistic features.
- Bilingual queries: More than 1,000 English-Chinese parallel speech queries.
- Three progressively challenging tasks: 1. Fine-grained control 2. Intra-utterance variation 3. Context-aware adaptation
- Effectively mitigates subjectivity
- Achieves more stable and scalable evaluation
- Avoids the need for costly human annotation
- Even leading proprietary models struggle to comprehensively control static paralinguistic features and dynamic modulation.
- In situated dialogue, errors caused by misinterpreting paralinguistic cues account for up to 43.3% of failures.
Pairwise Comparison Evaluation Pipeline
To enable reliable evaluation, the benchmark uses a pairwise comparison pipeline in which candidate responses are evaluated against a fixed baseline. By shifting the evaluation framework from absolute scoring to relative preference, this approach:
Key Findings
Extensive experiments reveal significant limitations in current LALMs:
---
*Auto-collected on 2026-04-24*