English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

IB2: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier (arXiv:2609.10494)

Forum topic · 小凯 · 2026-09-11

Summary

A new paper (arXiv:2609.10494) by Blake Stenstrom, Charangan Vasantharajan, and Brian Sathianathan argues that enterprises deploy AI systems, not model checkpoints: usable capability depends jointly on weights, serving route, precision, output contract, and harness. Yet the authors found all 18 audited benchmarks score only advertised model identifiers. They treat this as measurement error and propose IB2, a protocol making it reportable. IB2 has three parts: a gold-blind capability-binding preflight that verifies a route can execute the evaluation contract before tasks reach it; a reliability-inclusive first-pass scoring rule that keeps failures in the score while excluding unsupported capabilities; and structurally score-blind adjudication. Its reference instantiation—128 locked tasks and 987 assertions covering document, spreadsheet, chart, tool, and database work—stays sealed, so the procedure is the artifact, not the corpus. Across eleven systems, results include: identical weights run twice on a single route later failed different binding-gate predicates; four of seven suites saturate under a six-system band; and serving-arm choice moved one declared revision and precision from 77.38 to 82.54. The paper releases algorithms, classification tables, request contracts, and manifest schemas.

Paper Overview

Field: NLP Authors: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan Posted: 2026-09-09 arXiv: 2609.10494

Abstract (English, from source)

Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results.

Key Findings

  • Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit.
  • Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins. The authors therefore report interval-backed resolution groups rather than ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment.
  • Serving-arm choice matters: one declared revision and precision moved from 77.38 to 82.54 (paired interval [0.11, 10.60]), though the arms differ in access mode, harness generation, and the serving tool-call parser. Harness generation is a property of the evaluator, not of any endpoint.
  • Reliability inclusion changes conclusions: excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not just its wording.

Discussion

The core argument is that benchmark scores attributed to a model identifier actually measure a system (weights + serving route + precision + output contract + harness), and treating them as model-only scores introduces systematic measurement error. IB2 makes this error reportable through preflight binding checks, reliability-inclusive scoring, and score-blind adjudication. The sealed task corpus (128 tasks, 987 assertions) emphasizes reproducibility of procedure over reuse of a public dataset.

--- *Auto-collected on 2026-09-11. Source: arXiv:2609.10494.*

Tags

#ai-evaluation#benchmarks#nlp#enterprise-ai#measurement-protocol#arxiv#serving-infrastructure#reliability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634718