English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Forum topic · 小凯 · 2026-07-08

Summary

SPEARBench (arXiv:2607.05365) is a new benchmark for evaluating the naturalness of streaming speech-to-speech language models, which answer spoken queries directly with synthesized speech. Developed by Thomas Thebaud, Yuzhe Wang, Hao Zhang, and colleagues at Amazon, the benchmark builds controlled conversational prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates responses along a multi-dimensional protocol covering response latency, interruptions, speech quality, ASR robustness, language and dialect consistency, emotional naturalness, interpersonal posture, and interpretable distribution baselines. Findings show that although current models achieve high signal-level quality and low ASR error rates, their behavior still differs from human conversational dynamics in latency, overlapping speech, dialect preservation, emotional adaptation, and interpersonal posture dynamics, highlighting a gap between technical metrics and perceived conversational naturalness.

Overview

  • Field: NLP
  • Authors: Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez
  • Published: 2026-07-06
  • arXiv: 2607.05365
  • Abstract

    Streaming speech-to-speech language models aim to answer spoken queries directly with synthesized speech. However, standard speech and text benchmarks fail to capture whether these systems feel natural in conversation, where timing, turn-taking, prosody, interpersonal posture, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.

    This paper introduces SPEARBench, a benchmark for evaluating the naturalness of speech-to-speech language models in question-answering interactions. SPEARBench constructs controlled conversational prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates generated responses using a multi-dimensional protocol covering:

  • Response latency
  • Interruptions and overlapping speech
  • Speech quality
  • ASR robustness
  • Language and dialect consistency
  • Emotional naturalness
  • Interpersonal posture
  • Interpretable distribution baselines

Key Findings

Current models can achieve high signal-level quality and low ASR error rates, but they still diverge from human conversational behavior in terms of latency, overlapping speech, dialect preservation, emotional adaptation, and interpersonal posture dynamics.

---

*Auto-collected on 2026-07-06.*

Tags

#speech-to-speech#benchmark#nlp#speech-synthesis#conversational-ai#arxiv#speech-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346227