[论文] On the estimation and validity of AI time horizons---a statistical loo...
研究领域: ML 作者: Drew T. Nguyen, William Fithian 发布时间: 2026-10-08 arXiv: 2610.12466
论文概要
研究领域: ML 作者: Drew T. Nguyen, William Fithian 发布时间: 2026-10-08 arXiv: 2610.12466
中文摘要
METR 的 50% 时间跨度衡量 AI 以 50% 概率解决的软件任务对应的人类完成时间,使 AI 能力可用可解释的单位表达。在 228 个任务和 26 个 AI 系统上,我们用样条函数和项目反应理论重新计算时间跨度,放松了"任务 AI 难度依赖于人类时间对数的线性函数"这一假设。拟合的样条可解释为将人类时间"转换"为 AI 难度的函数:在 2-30 分钟区间几乎平坦,其他区域接近线性。因此从 3 分钟到 30 分钟的跳跃远比从 30 分钟到 5 小时容易,尽管倍数同为 10 倍。我们的贡献包括:在交叉验证的严格评分规则下表现更优的时间跨度点估计,以及评估构念效度的诊断图。建议将时间跨度与诊断图结合解读,尤其在新基准提出或现有基准扩展至更长任务时。
原文摘要
METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of \(10 \times\). Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of ...
*自动采集于 2026-10-10*
#论文 #arXiv #ML #小凯