[论文] IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier (arXiv:2609.10494)
论文概要 研究领域: NLP 作者: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan 发布时间: 2026-09-09 arXiv: 2609.10494
论文概要
研究领域: NLP 作者: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan 发布时间: 2026-09-09 arXiv: 2609.10494中文摘要
企业部署的是系统,不是检查点。可用能力 jointly 取决于权重、服务路由、精度、输出契约和 harness,但所有18个审计基准都只给广告中的模型标识符打分。作者将这种系统性偏差视为测量误差,并提出 IB2 协议使其可报告。协议包含三部分:gold-blind 能力绑定预检,在任务到达前验证路由能执行评估契约;可靠性包含的首次评分规则,将失败保留在分数中同时将不支持的能力排除;以及结构上分数盲的裁决。在11个系统上的测试发现:能力可用性是可测量的——相同权重上的两次完整单路由运行后来在不同绑定门谓词上失败;区分度不均匀——七个套件中有四个在六系统带宽下饱和;服务路由选择使一个声明版本和精度从77.38变为82.54。原文摘要
Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.*自动采集于 2026-09-11*
#论文 #arXiv #NLP #小凯