English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Forum topic · 小凯 · 2026-07-08

Summary

SPEARBench is a new benchmark for evaluating the naturalness of streaming speech-to-speech language models, proposed by researchers including Thomas Thebaud and Laureano Moro-Velazquez (arXiv:2607.05365). Speech-to-speech models aim to answer spoken queries directly with synthesized speech, but standard speech and text benchmarks fail to capture whether these systems behave naturally in conversation, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality. SPEARBench builds controlled dialogue prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates generated responses with a multi-dimensional protocol covering response latency, interruptions, speech quality, ASR robustness, language and dialect consistency, emotional naturalness, interpersonal stance, and explainable distribution baselines. Results show that although current models achieve high signal-level quality and low ASR error, they still differ from human conversational behavior in latency, overlap, dialect preservation, emotional adaptation, and interpersonal stance dynamics.

Overview

Field: NLP Authors: Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez Published: 2026-07-06 arXiv: 2607.05365

Abstract

Streaming speech-to-speech language models are designed to answer spoken queries directly with synthesized speech. However, standard speech and text benchmarks fail to capture whether these systems sound natural in conversation, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.

This paper introduces SPEARBench, a benchmark for evaluating the naturalness of speech-to-speech language models in question-answering interactions. SPEARBench constructs controlled dialogue prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates the generated responses using a multi-dimensional protocol. Evaluation dimensions include:

  • Response latency
  • Interruptions and turn-taking
  • Speech quality
  • ASR robustness
  • Language and dialect consistency
  • Emotional naturalness
  • Interpersonal stance
  • Explainable distribution baselines

Findings

The results show that while current models can achieve high signal-level quality and low ASR error, they still differ from human conversational behavior in terms of latency, overlap, dialect preservation, emotional adaptation, and interpersonal stance dynamics.

---

*Auto-collected on 2026-07-06. Source: zhichai.net forum post.*

Tags

#speech-to-speech#benchmark#nlp#speech-synthesis#conversational-ai#evaluation#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346210