[论文] Available Guardrails: Certifying Selective Prediction across ML System...
研究领域: ML 作者: Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky 发布时间: 2026-09-18 arXiv: 2609.22048
论文概要
研究领域: ML 作者: Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky 发布时间: 2026-09-18 arXiv: 2609.22048
中文摘要
选择性预测器充当安全门:仅当预测显得足够可信时才返回输出。越来越多的部署要求在目标精度下,对每个相关报告单元(如工具、策略标签或患者亚组)的可靠性进行认证。主要困难往往不在于已颁发的证书是否有效,而在于有限的校准数据能否产生证书:当门控变得更安全或更细粒度时,某些单元可能因证据太少而无法认证。我们通过经典的精确二项分布反演使“可用性”这一概念可计算,并在固定组顺序下将报告划分选择表述为一个动态规划,揭示安全性、粒度和承载流量之间的权衡。所得前沿揭示出一个被有限样本估计几乎抹平的大量人口机会:一个“知情真实”的规划者比支持度平衡多获得 0.157 的平均覆盖率,而朴素估计器只能恢复 0.005——这使从有限数据中恢复成为核心挑战。在一个规划划分上构建候选划分、在另一个划分上进行选择,可部分弥合这一差距,使平均覆盖率比支持度平衡提高 0.060,且在三个意图路由数据集和两种架构的 60 个模型效应中有 59 个复现了这一方向。一个互补的保持有效性的杠杆——在报告单元间重新分配族系错误预算——在人口量与噪声估计下都能恢复额外覆盖率。同一前沿以预测器特定的上限反复出现在 LLM 工具调用、内容审核、病灶分类和推荐等场景中。因此,认证可用性是一种可规划的部署资源,决定了安全门何时能被认证、以何种粒度、以及覆盖多少流量。
原文摘要
A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The r...
*自动采集于 2026-09-22*
#论文 #arXiv #ML #小凯