Overview
- Field: NLP
- Authors: Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez
- Published: 2026-07-06
- arXiv: 2607.05365
- Response latency
- Interruptions and overlapping speech
- Speech quality
- ASR robustness
- Language and dialect consistency
- Emotional naturalness
- Interpersonal posture
- Interpretable distribution baselines
Abstract
Streaming speech-to-speech language models aim to answer spoken queries directly with synthesized speech. However, standard speech and text benchmarks fail to capture whether these systems feel natural in conversation, where timing, turn-taking, prosody, interpersonal posture, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.
This paper introduces SPEARBench, a benchmark for evaluating the naturalness of speech-to-speech language models in question-answering interactions. SPEARBench constructs controlled conversational prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates generated responses using a multi-dimensional protocol covering:
Key Findings
Current models can achieve high signal-level quality and low ASR error rates, but they still diverge from human conversational behavior in terms of latency, overlapping speech, dialect preservation, emotional adaptation, and interpersonal posture dynamics.
---
*Auto-collected on 2026-07-06.*