Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Differentiator
Field: Machine Learning Author: George Andrikopoulos Published: 2026-08-19 arXiv: 2608.19140
Key points
- Frontier language models are compared, marketed, and benchmarked on *capability* — what their best or average output can achieve. The author argues this measures the wrong axis.
- Models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision — how tightly outputs concentrate around that target across repeated, identical requests.
- Borrowing the marksman's distinction: capability is where the average shot lands; reliability is the size of the group.
- *Consistent failures*: tight groups off center, correctable through operational discipline (aim adjustment).
- *Scattered failures*: wide groups, correctable only by changing the model or its sampling (the rifle problem).
- A defined grouping metric and a specified tooling workflow.
- Tracking the grouping of human–AI pairs over time yields a compound signal needed for the field study in paper 1.
Three claims
1. Precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. 2. Precision is measurable cheaply and without circularity: run a fixed suite of deterministically scoreable tasks at fixed temperature multiple times and compute consistency of per-task results — no model-in-the-loop scorer needed. 3. The measurement is decision-guiding, not merely descriptive. It distinguishes:
Contributions
First field run
The initial real-world run (since replicated) illustrates both the method and its most important limitation: a measured gap was fully closed by a single rule (0/5 → 5/5), while a task suite authored from that rule itself found no value — because frontier models already embody the explicit good practices. This establishes that a discipline's value is discovered by measurement on real work, not by constructing tasks from its own rulebook.
--- *Source: arXiv:2608.19140*