English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

Forum topic · 小凯 · 2026-05-21

Summary

DecisionBench is a benchmark substrate for studying emergent delegation in long-horizon agentic workflows, released on arXiv (2505.01259) by Yuxuan Gao, Megan Wang, and Yi Ling Yu. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a peer-model pool of 11 models from 7 vendor families, a delegation interface (call_model plus optional read_profile), a deterministic skill-annotation layer, and metrics covering quality, cost, latency, delegation rate, routing fidelity-at-k, vendor self-preference, and a counterfactual-delegation ceiling. Because the substrate is agnostic to how peer information is generated or delivered, learned routers, peer memories, adaptive profile construction, and multi-step delegation can all be evaluated against it. A five-condition reference sweep (n=23,375 task instances) shows: (i) mean end-task quality is statistically indistinguishable across four awareness conditions (|beta| <= 0.010, p >= 0.21); (ii) routing fidelity-at-1 varies from 7.5% to 29.5% at near-equal quality, with delivery channel dominating over description content; (iii) a counterfactual ceiling places perfect delegation 15-31 percentage points above measured performance. The substrate, annotation layer, intervention suite, analysis pipeline, and 220 run archives are publicly released.

Paper Overview

Research areas: cs.AI, cs.CL, cs.MA Authors: Yuxuan Gao, Megan Wang, Yi Ling Yu Published: 2026-05-21 arXiv: 2505.01259

Abstract

We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau-bench, BFCL multi-turn), a peer-model pool (11 models, 7 vendor families), a delegation interface (call_model plus an optional read_profile channel), a deterministic skill-annotation layer, and a multi-axis metric suite covering quality, cost, latency, delegation rate, routing fidelity-at-k, vendor self-preference, and a counterfactual-delegation ceiling. The substrate is agnostic to how peer information is generated or delivered, so learned routers, richer peer memories, adaptive profile construction, and multi-step delegation can all be evaluated against it.

Key Findings

From a five-condition reference sweep on the full pool (n=23,375 task instances):

1. Quality alone misses the orchestration signal — mean end-task quality is statistically indistinguishable across the four awareness conditions (|beta| <= 0.010, p >= 0.21). 2. Routing fidelity varies widely at equal quality — routing fidelity-at-1 ranges from 7.5% to 29.5% across conditions at near-equal mean quality; the delivery channel (on-demand tool vs. preloaded description) dominates description content. 3. Large unrealized headroom — a counterfactual ceiling places perfect delegation 15-31 percentage points above measured performance on every suite, motivating future orchestration methods.

The authors release the substrate, annotation layer, reference intervention suite, analysis pipeline, and 220 per-condition run archives.

--- *Auto-collected on 2026-05-21*

Tags

#ai-agents#benchmark#llm-orchestration#delegation#routing#multi-agent-systems#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620526