English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DuplexSpeechBench-IFEval: A Benchmark for Implicit Instruction Following in Full-Duplex Voice Agents

Forum topic · 小凯 · 2026-09-06

Summary

DuplexSpeechBench-IFEval (DSB-IFEval) is a new benchmark from researchers at the University of Maryland (Puneet Mathur and Dinesh Manocha, arXiv:2509.00008) that evaluates how well full-duplex voice agents follow implicit instructions. Unlike prior benchmarks that rely on explicit turn-management commands, DSB-IFEval tests whether agents can infer appropriate conversational behavior—listening, backchanneling, interrupting, handling overlaps, and taking or yielding the floor—from a persona alone. The benchmark includes 1,038 test cases across eight assistant roles and five conditioning protocols: default behavior, explicit instructions, persona-implied behavior, combined persona-rule conditioning, and instruction conflict. Evaluation uses a deterministic Instruction Adherence Score (IAS) for turn management and an LLM-judged Persona Adherence Score (PAS) for persona-consistent content. Testing six real-time speech systems reveals architectural trade-offs: full-duplex models like F-Actor and PersonaPlex drop 9.7% and 4.5% in adherence when behavior must be inferred from persona, while GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly follow persona-consistent content but show rigid turn-taking behavior and limited proactive actions. All systems struggle to override instructions under safety conflicts.

Overview

Field: AI/ML Authors: Puneet Mathur, Dinesh Manocha Published: 2026-09-06 arXiv: 2509.00008

Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. DuplexSpeechBench-IFEval (DSB-IFEval) was introduced to evaluate implicit instruction-following in real-time spoken interaction.

Benchmark Design

  • 1,038 test cases spanning eight diverse assistant roles
  • Five conditioning protocols for instruction-following:
  • 1. Default behavior 2. Explicit behavioral instructions 3. Persona-implied behavior 4. Combined persona-rule conditioning 5. Instruction conflict
  • Metrics:
  • Deterministic Instruction Adherence Score (IAS) for real-time floor management
  • LLM-judged Persona Adherence Score (PAS) for persona-consistent content
  • Key Findings

    Across six real-time speech systems, the authors found architecture-dependent trade-offs:

  • Full-duplex models such as F-Actor and PersonaPlex are sensitive to whether conversational behavior is explicitly stated versus inferred from a persona, with adherence dropping 9.7% and 4.5% respectively under persona-only conditions.
  • GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly follow persona-consistent content, but their turn-taking behavior does not adapt between explicit and persona-only instructions, and they remain restricted on several proactive actions.
  • Even when systems reliably follow instructions that conflict with their assigned persona, they struggle to override these instructions under safety conflicts.

Conclusion

The results show that inferring role-implied behavior, executing it at appropriate conversational moments, and resolving competing instructions remain distinct challenges for full-duplex voice agents.

Source: arXiv:2509.00008

Tags

#full-duplex-speech#voice-agents#benchmark#instruction-following#speech-interaction#llm-evaluation#ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634532