English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Differentiator for LLMs

Forum topic · 小凯 · 2026-08-21

Summary

This arXiv paper (2608.19140) by George Andrikopoulos argues that frontier language models should be evaluated on precision rather than capability. Borrowing the marksman's distinction between where the average shot lands and the size of the shot group, the author claims that models have saturated accuracy: their mean outputs land on target, so what now differentiates systems is how tightly outputs cluster around that target across repeated, identical requests. Three claims are made: benchmark culture systematically reports central tendency instead of spread and thus misses precision; precision can be measured cheaply and without circularity by running a fixed suite of deterministically scoreable tasks at fixed temperature and computing per-task consistency, requiring no model-in-the-loop scoring; and the measurement is decision-guiding, distinguishing consistent failures (tight groups off center, correctable via operational discipline) from scattered failures (wide groups, requiring model or sampling changes). The paper defines a grouping metric, specifies a tool, and describes a first field run showing a measured gap fully closed by a single rule (0/5 to 5/5), while a task suite derived from the rule itself found no value because frontier models already embody explicit good practices.

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Differentiator

Field: Machine Learning Author: George Andrikopoulos Published: 2026-08-19 arXiv: 2608.19140

Key points

  • Frontier language models are compared, marketed, and benchmarked on *capability* — what their best or average output can achieve. The author argues this measures the wrong axis.
  • Models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision — how tightly outputs concentrate around that target across repeated, identical requests.
  • Borrowing the marksman's distinction: capability is where the average shot lands; reliability is the size of the group.
  • Three claims

    1. Precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. 2. Precision is measurable cheaply and without circularity: run a fixed suite of deterministically scoreable tasks at fixed temperature multiple times and compute consistency of per-task results — no model-in-the-loop scorer needed. 3. The measurement is decision-guiding, not merely descriptive. It distinguishes:

  • *Consistent failures*: tight groups off center, correctable through operational discipline (aim adjustment).
  • *Scattered failures*: wide groups, correctable only by changing the model or its sampling (the rifle problem).
  • Contributions

  • A defined grouping metric and a specified tooling workflow.
  • Tracking the grouping of human–AI pairs over time yields a compound signal needed for the field study in paper 1.

First field run

The initial real-world run (since replicated) illustrates both the method and its most important limitation: a measured gap was fully closed by a single rule (0/5 → 5/5), while a task suite authored from that rule itself found no value — because frontier models already embody the explicit good practices. This establishes that a discipline's value is discovered by measurement on real work, not by constructing tasks from its own rulebook.

--- *Source: arXiv:2608.19140*

Tags

#llm-evaluation#precision-vs-capability#benchmarks#reliability#model-consistency#machine-learning#arxiv#evaluation-methodology

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633745