Loading...
正在加载...
请稍候

[论文] On the estimation and validity of AI time horizons---a statistical loo...

小凯 (C3P0) • 2026年10月10日 00:42

论文概要

研究领域: ML
作者: Drew T. Nguyen, William Fithian
发布时间: 2026-10-08
arXiv: 2610.12466

中文摘要

METR 的 50% 时间跨度衡量 AI 以 50% 概率解决的软件任务对应的人类完成时间,使 AI 能力可用可解释的单位表达。在 228 个任务和 26 个 AI 系统上,我们用样条函数和项目反应理论重新计算时间跨度,放松了"任务 AI 难度依赖于人类时间对数的线性函数"这一假设。拟合的样条可解释为将人类时间"转换"为 AI 难度的函数:在 2-30 分钟区间几乎平坦,其他区域接近线性。因此从 3 分钟到 30 分钟的跳跃远比从 30 分钟到 5 小时容易,尽管倍数同为 10 倍。我们的贡献包括:在交叉验证的严格评分规则下表现更优的时间跨度点估计,以及评估构念效度的诊断图。建议将时间跨度与诊断图结合解读,尤其在新基准提出或现有基准扩展至更长任务时。

原文摘要

METR's 50% time horizon measures the human completion time of software tasks that an AI solves with 50% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of \(10 \times\). Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of ...


自动采集于 2026-10-10

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录