English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

Forum topic · 小凯 · 2026-09-06

Summary

DuplexSpeechBench-IFEval (DSB-IFEval), introduced by Puneet Mathur and Dinesh Manocha (arXiv:2509.00008), is a benchmark for evaluating implicit instruction-following in real-time full-duplex voice agents. While existing benchmarks test turn-management via explicit instructions, deployed agents are usually configured through personas from which conversational behavior—listening, backchanneling, interrupting, handling overlaps, taking and yielding the floor—must be inferred. DSB-IFEval contains 1,038 test cases across eight assistant roles and five conditioning protocols: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona-rule conditioning, and instruction conflict. Evaluation uses a deterministic Instruction Adherence Score (IAS) for turn management and an LLM-judged Persona Adherence Score (PAS) for persona-consistent content. Across six real-time speech systems, full-duplex models like F-Actor and PersonaPlex proved sensitive to whether behavior is stated or inferred, dropping 9.7% and 4.5% under persona-only conditions, while GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat followed persona content strongly but adapted turn-taking behavior poorly and remained limited on proactive actions. Systems also struggle to override persona instructions under safety conflicts.

Overview

  • Field: AI/ML
  • Authors: Puneet Mathur, Dinesh Manocha
  • Published: 2026-09-06
  • arXiv: 2509.00008
  • Abstract

    Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. The authors introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction.

    Benchmark Design

  • 1,038 test cases spanning eight diverse assistant roles
  • Five conditioning protocols for instruction-following:
  • 1. Default behavior 2. Explicit behavioral instructions 3. Persona-implied behavior 4. Combined persona–rule conditioning 5. Instruction conflict
  • Metrics:
  • Deterministic Instruction Adherence Score (IAS) for real-time floor management
  • LLM-judged Persona Adherence Score (PAS) for persona-consistent content
  • Key Findings

    Across six real-time speech systems, the evaluation reveals architecture-dependent trade-offs:

  • Full-duplex models such as F-Actor and PersonaPlex are more sensitive to whether conversational behavior is explicitly stated versus inferred from a persona, with adherence dropping by 9.7% and 4.5% respectively under persona-only conditions.
  • GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly follow persona-consistent content, but their floor-management behavior does not adapt between explicit and persona-only instructions, and they remain restricted on several proactive actions.
  • Even when systems reliably follow instructions that conflict with their assigned persona, they struggle to override these instructions under safety conflicts.

Conclusion

Inferring persona-implied behavior, executing it at appropriate conversational moments, and resolving competing instructions remain distinct challenges for full-duplex voice agents.

---

*Source: arXiv:2509.00008*

Tags

#full-duplex-speech#voice-agents#benchmark#instruction-following#persona#turn-taking#llm-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634522