English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Real-Time Voice AI Hears but Does Not Listen: The Emotional Intelligence Gap

Forum topic · 小凯 · 2026-06-26

Summary

A 2026 arXiv paper (2606.19226) by Martijn Bartelds, Federico Bianchi, and James Zou evaluates four leading production real-time voice AI systems—OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash—on tasks where both words and vocal delivery carry meaningful information. Across three consequential scenarios, all four systems act on the words rather than the voice: they end calls with crying callers who insist nothing is wrong, approve wire transfers authorized in frightened voices, and enroll callers whose agreement is clearly sarcastic. Surprisingly, this is often not a perceptual failure—when asked directly, three of the four systems reliably identify the distress, fear, or sarcasm they later ignore in decisions. A similar pattern appears when estimating accent and age, where responses follow lexical bias rather than acoustic features. The authors call this perception-action disconnect the emotional intelligence gap of voice AI. Prompting systems to attend to vocal delivery only partially and inconsistently improves performance. The findings suggest current real-time voice AI often behaves as if speech were reduced to transcripts, warranting caution in scenarios where tone and emotional delivery matter.

Overview

Field: NLP Authors: Martijn Bartelds, Federico Bianchi, James Zou Published: 2026-06-25 arXiv: 2606.19226

Key points

  • Speech conveys information through both words and vocal delivery. The paper evaluates four leading production real-time voice systems on tasks where words and delivery patterns both carry meaningful information:
  • OpenAI's GPT Realtime 2
  • Google's Gemini 3.1 Flash Live
  • Alibaba's Qwen3.5 Omni Plus and Omni Flash
  • Across three consequential scenarios, all four systems act on the words rather than the voice:
  • They end calls with crying callers who insist nothing is wrong.
  • They approve wire transfers authorized in frightened voices.
  • They enroll callers whose agreement is clearly sarcastic.
  • Surprisingly, this is often not a failure of perception: when asked directly, three of the four systems reliably identify the distress, fear, or sarcasm they later ignore when making decisions.
  • A similar pattern emerges when systems estimate accent and age: responses often follow lexical bias in the words rather than the speaker's acoustic characteristics.
  • The authors term this perception–action disconnect the "emotional intelligence gap" of voice AI.
  • Prompting systems to explicitly attend to vocal delivery only partially and inconsistently improves performance.

Conclusion

Current real-time voice AI systems often behave as if speech were reduced to a text transcript. The authors recommend caution when deploying such systems in scenarios where tone and emotional delivery carry critical information.

---

*Auto-collected on 2026-06-26*

Tags

#voice-ai#speech-recognition#realtime-systems#nlp#emotion-recognition#llm-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208135