论文概要
研究领域: ML
作者: Sho Kawano, Zehang Richard Li, Paul A. Parker
发布时间: 2026-09-17
arXiv: 2609.20758
中文摘要
评估 AI 系统需要做细分(disaggregated)评估,因为性能会随领域而变化——例如基准任务类型或部署 Agent 中的对话类型。穷尽式测试成本高昂,评估只能基于有标注单元的抽样。我们把评估集视为有限总体,目标是给出每个领域均值的准确点估计与区间估计。直接使用估计量(包括预测驱动推断 PPI)只利用该领域自身的标注,在标注稀少的领域不精确。小区域估计(small area estimation)正是针对此问题,我们在其基础上构建了估计与验证一体化工作流。估计方面,提出预测驱动平滑(PP-S):对每个领域的 PPI 估计拟合贝叶斯模型,并扩展出跨报告分类体系借力信息的形式(PP-TS)。验证方面,导出一个新的近似无偏的设计型交叉验证分数,用于在直接估计量与平滑估计量之间做选择。我们在带可验证评分的精选基准和人工评分的部署 Agent 流量上开展研究——两者每个结局都被完整观测。结果显示:所提估计量在点估计与区间估计上均优于直接估计量,覆盖率接近名义水平;在相同抽样预算下,我们的选择分数的表现与独立验证集相当,且对所选估计量误差的估计准确得多。
原文摘要
Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension that borrows str...
自动采集于 2026-09-20
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。