English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation

Forum topic · 小凯 · 2026-04-24

Summary

SpeechParaling-Bench is a new benchmark for evaluating paralinguistic awareness in Large Audio-Language Models (LALMs), addressing coarse feature coverage and subjective assessment in prior evaluations. The benchmark expands coverage from fewer than 50 to over 100 fine-grained paralinguistic features and includes more than 1,000 English-Chinese parallel speech queries organized into three progressively difficult tasks: fine-grained control, intra-utterance variation, and context-aware adaptation. For reliable evaluation, the authors develop a pairwise comparison pipeline in which candidate responses are judged against a fixed baseline, shifting from absolute scoring to relative preference. This reduces subjectivity and enables stable, scalable assessment without costly human annotation. Experiments reveal significant limitations in current LALMs: even leading proprietary models struggle to fully control static paralinguistic features and dynamic modulation, and up to 43.3% of errors in situated dialogue stem from misinterpreting paralinguistic cues. The findings underscore the need for more robust paralinguistic modeling to advance human-aligned speech assistants. Paper: arXiv 2604.20842.

Paper Overview

Field: NLP Authors: Ruohan Liu, Shukang Yin, Tao Wang Released: 2026-04-22 arXiv: 2604.20842

Introduction

Paralinguistic cues—the non-verbal aspects of speech such as tone, emotion, and prosody—are essential for natural human-computer interaction. However, evaluating these capabilities in Large Audio-Language Models (LALMs) has been limited by coarse feature coverage and the inherent subjectivity of assessment methods.

SpeechParaling-Bench

To address these challenges, the authors introduce SpeechParaling-Bench, a comprehensive benchmark for paralinguistic-aware speech generation with the following features:

  • Expanded feature coverage: Grows from fewer than 50 to over 100 fine-grained paralinguistic features.
  • Bilingual queries: More than 1,000 English-Chinese parallel speech queries.
  • Three progressively challenging tasks:
  • 1. Fine-grained control 2. Intra-utterance variation 3. Context-aware adaptation

    Pairwise Comparison Evaluation Pipeline

    To enable reliable evaluation, the benchmark uses a pairwise comparison pipeline in which candidate responses are evaluated against a fixed baseline. By shifting the evaluation framework from absolute scoring to relative preference, this approach:

  • Effectively mitigates subjectivity
  • Achieves more stable and scalable evaluation
  • Avoids the need for costly human annotation
  • Key Findings

    Extensive experiments reveal significant limitations in current LALMs:

  • Even leading proprietary models struggle to comprehensively control static paralinguistic features and dynamic modulation.
  • In situated dialogue, errors caused by misinterpreting paralinguistic cues account for up to 43.3% of failures.
These findings highlight the urgent need for more robust paralinguistic modeling to advance speech assistants toward human-aligned behavior.

---

*Auto-collected on 2026-04-24*

Tags

#speech-paralinguistics#audio-language-models#benchmark#speech-generation#nlp#evaluation#arxiv#lalms

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618681