论文概要
研究领域: ML
作者: George Andrikopoulos
发布时间: 2026-08-19
arXiv: 2608.19140
中文摘要
前沿语言模型在能力上进行比较、营销和基准测试——它们的最佳或平均输出能实现什么。我认为这测量了错误的维度。模型已经饱和了准确性:它们的平均输出落在目标上。现在在实践中区分一个系统与另一个系统的是精度:在重复、相同的请求中,它们的输出围绕该目标的集中程度。借用射手的区分,能力是平均射击落在哪里;可靠性是组的大小。我提出三个主张。首先,精度而非能力是系统之间的前沿区分因素,基准测试文化系统性地未能测量它,报告中心趋势而非离散程度。其次,精度可以通过在固定温度下多次运行固定套件的确定性评分任务并计算每任务结果的一致性来廉价且无循环地测量——无需模型在环评分器。第三,测量不仅仅是描述性的,而且是决策指导性的:它区分一致的失败(偏离中心的紧组,可通过操作纪律纠正——瞄准调整)和分散的失败(宽组,只能通过更改模型或其采样来纠正——步枪问题)。我定义了一个分组指标,指定了一个工具,并展示跟踪人-AI对随时间的分组如何产生论文1的实地研究所需的复合信号。第一次真实运行,自那以后被复制,说明了该方法及其最重要的限制:一个测量的差距被单一规则完全关闭(0/5 -> 5/5),而从规则本身编写的一套任务没有发现价值,因为前沿模型已经体现了明确的良好实践——确立了学科的价值通过在真实工作上的测量来发现,而不是从其自己的规则书中构建。
原文摘要
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circulari...
自动采集于 2026-08-21
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。