English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Forum topic · 小凯 · 2026-09-20

Summary

This arXiv paper (2609.20758) by Sho Kawano, Zehang Richard Li, and Paul A. Parker addresses disaggregated AI system evaluation, where performance varies across domains such as benchmark task types or conversation types in deployed agents. Treating the evaluation set as a finite population, the authors seek accurate point and interval estimates of each domain mean when labels are sampled sparsely. Direct estimators like prediction-powered inference (PPI) rely only on a domain's own labels and become imprecise in low-label domains. Building on small area estimation, they propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's PPI estimate, plus an extension (PP-TS) that borrows information across reporting taxonomies. For validation, they derive a novel approximately unbiased design-based cross-validation score to choose between direct and smoothed estimators. Experiments on curated benchmarks with verifiable scores and human-scored deployed agent traffic show the proposed estimators beat direct estimators with near-nominal coverage, and their selection score matches independent validation sets at the same sampling budget.

Overview

Field: Machine Learning Authors: Sho Kawano, Zehang Richard Li, Paul A. Parker Published: 2026-09-17 arXiv: 2609.20758

Summary

Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. The paper treats the evaluation set as a finite population and seeks accurate point and interval estimates of each domain mean.

Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and the authors build on it to develop an integrated workflow for estimation and validation:

  • Estimation: They propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension (PP-TS) that borrows strength across reporting taxonomies.
  • Validation: They derive a new approximately unbiased design-based cross-validation score for choosing between direct and smoothed estimators.
Experiments are conducted on curated benchmarks with verifiable scores and on deployed agent traffic with human scoring—in both cases with fully observed outcomes. Results show that the proposed estimators outperform direct estimators in both point and interval estimation, with coverage close to nominal levels. Under the same sampling budget, their selection score performs comparably to an independent validation set and estimates the error of the selected estimator far more accurately.

--- *Auto-collected 2026-09-20*

Tags

#machine-learning#arxiv#prediction-powered-inference#small-area-estimation#ai-evaluation#bayesian-modeling#cross-validation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635003