English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

Forum topic · 小凯 · 2026-09-11

Summary

Enterprises deploy AI systems, not model checkpoints: usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet audited benchmarks score only advertised model identifiers. The IBIB (IB2) protocol treats this as measurement error and makes it reportable through three components: a gold-blind capability-binding preflight that verifies a route can execute the evaluation contract before tasks arrive, a reliability-inclusive first-pass scoring rule that keeps failures in scores while excluding unsupported capabilities, and structurally score-blind adjudication. The reference instantiation uses 128 locked tasks and 987 assertions spanning document, spreadsheet, chart, tool, and database work, with the procedure rather than the corpus treated as the artifact. Testing across eleven systems showed capability availability is measurable (two identical-weight single-route runs failed different binding-gate predicates), discrimination is uneven (four of seven suites saturate, reported as interval-backed resolution groups rather than ranks), and serving-arm choice moved one declared revision and precision from 77.38 to 82.54. Reliability inclusion changed a conclusion, not merely its wording. Source: arXiv 2609.10494.

Paper Overview

  • Field: NLP
  • Authors: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
  • Published: 2026-09-09
  • arXiv: 2609.10494
  • Key Points

    Enterprises deploy *systems*, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness — yet all 18 audited benchmarks score advertised model identifiers. The authors treat this systematic bias as measurement error and propose the IB2 protocol to make it reportable.

    The protocol has three parts:

  • A gold-blind capability-binding preflight that verifies a route can execute the evaluation contract before any task reaches it.
  • A reliability-inclusive first-pass scoring rule that keeps failure in the score while keeping unsupported capability out.
  • Structurally score-blind adjudication.
Its reference instantiation — 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool, and database work — stays sealed: the procedure is the artifact, not the corpus.

Findings Across Eleven Systems

1. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. 2. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so the authors report interval-backed resolution groups rather than ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. 3. Serving-arm choice matters: it moved one declared revision and precision from 77.38 to 82.54 (paired interval [0.11, 10.60]), though the arms differ in access mode, harness generation, and the serving tool-call parser — and harness generation is a property of the evaluator, not any endpoint. 4. Reliability inclusion changes conclusions: excluding failed responses from denominators changes the point ordering.

The algorithms, classification tables, request contract, and manifest schemas are released with the paper.

---

*Auto-collected on 2026-09-11*

Tags

#papers#arxiv#nlp#ai-benchmarking#enterprise-ai#evaluation-protocol#measurement

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634709