Overview
Field: NLP Authors: Martijn Bartelds, Federico Bianchi, James Zou Published: 2026-06-25 arXiv: 2606.19226
Key points
- Speech conveys information through both words and vocal delivery. The paper evaluates four leading production real-time voice systems on tasks where words and delivery patterns both carry meaningful information:
- OpenAI's GPT Realtime 2
- Google's Gemini 3.1 Flash Live
- Alibaba's Qwen3.5 Omni Plus and Omni Flash
- Across three consequential scenarios, all four systems act on the words rather than the voice:
- They end calls with crying callers who insist nothing is wrong.
- They approve wire transfers authorized in frightened voices.
- They enroll callers whose agreement is clearly sarcastic.
- Surprisingly, this is often not a failure of perception: when asked directly, three of the four systems reliably identify the distress, fear, or sarcasm they later ignore when making decisions.
- A similar pattern emerges when systems estimate accent and age: responses often follow lexical bias in the words rather than the speaker's acoustic characteristics.
- The authors term this perception–action disconnect the "emotional intelligence gap" of voice AI.
- Prompting systems to explicitly attend to vocal delivery only partially and inconsistently improves performance.
Conclusion
Current real-time voice AI systems often behave as if speech were reduced to a text transcript. The authors recommend caution when deploying such systems in scenarios where tone and emotional delivery carry critical information.
---
*Auto-collected on 2026-06-26*