Overview
Field: NLP Authors: Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez Published: 2026-07-06 arXiv: 2607.05365
Abstract
Streaming speech-to-speech language models are designed to answer spoken queries directly with synthesized speech. However, standard speech and text benchmarks fail to capture whether these systems sound natural in conversation, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.
This paper introduces SPEARBench, a benchmark for evaluating the naturalness of speech-to-speech language models in question-answering interactions. SPEARBench constructs controlled dialogue prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates the generated responses using a multi-dimensional protocol. Evaluation dimensions include:
- Response latency
- Interruptions and turn-taking
- Speech quality
- ASR robustness
- Language and dialect consistency
- Emotional naturalness
- Interpersonal stance
- Explainable distribution baselines
Findings
The results show that while current models can achieve high signal-level quality and low ASR error, they still differ from human conversational behavior in terms of latency, overlap, dialect preservation, emotional adaptation, and interpersonal stance dynamics.
---
*Auto-collected on 2026-07-06. Source: zhichai.net forum post.*