What do you do when a thermometer breaks? Replace it. On September 4, Artificial Analysis released Intelligence Index v4.2, and that's exactly what it did: GPQA Diamond, the most famous public benchmark, was formally removed from the index. The official wording was brief — "an exceptional scientific reasoning evaluation that has now been saturated." Worn through. A ruler used for three years to measure "PhD-level scientific reasoning" has run out of scale.
To be clear what "saturated" means. GPQA is a PhD-level science multiple-choice exam, once considered among the hardest public benchmarks. When every frontier model climbs to 80, 90, 95 points, with scores crammed at the top and first place separated from fifth by two questions, the exam no longer measures intelligence — only who is better at that particular test.
What replaced the ruler
v4.2's answer is shifting weight toward tests models can't touch. Private, held-out test sets now account for 40% of the index — officially described as double the v4.1 level, i.e., up from 20%. The private set now includes three components: AA's own AA-Briefcase (multi-week real knowledge-work projects with thousands of source files, scored via rubrics plus pairwise comparison), AA-Omniscience, and reserved solutions from CritPt. The site also included a line worth underlining: "The held-out percentage will increase further in Index v5." The private share will keep rising.
The most interesting new private exam is called GDP.pdf, from Surge AI. The scale: 100 PDFs spanning ten domains, with evidence scattered across 4,592 pages of body text, tables, charts, footnotes, and red herrings. Each question is scored against 1,275 expert-written atomic criteria using a metric called All-pass Rate — one unmet criterion and the entire question scores zero. "All-or-nothing" is the sole standard.
The scorecard is chewy. On GDP.pdf, GPT-6 Astra scores 33.2%, GPT-5.6 Sol 28.2%, while the overall index leader Claude Fable 5.1 manages only 26.2%, in third place. On the main index, the order is Fable 5.1 first, Astra second (4 points above Sol), with Meta's lab third. On this private exam, the entire ranking flips.
The people swapping rulers are also being watched
In a 155-point Hacker News thread, the sharpest critique came from a user named redox99: in the previous index version, Astra and Sol were tied, and after being mocked the leaderboard "rushed to update the index so it fits what people expect" — scores were adjusted to match expectations. Another commenter asked directly: has this process been peer-reviewed? What are the sample sizes? That question has no good answer: AA is a commercial evaluation outfit that publishes no papers and submits to no review.
This site previously discussed Astra's 99.9% ARC ruler, concluding the publisher had "optimized the exam environment" for the benchmark. What AA did this time is the other side of the same coin: when model companies wear through public questions, evaluation shops hide the questions. Both point the same direction — [judgment] the public benchmark's role as industry referee is being actively deprecated. The rationale is sound; the cost is that evaluation transparency has retreated from "publicly verifiable" to "trust us."
There is no clean way out. Public rulers get worn through, private rulers can't be audited, and nobody is building the third path yet. The private weighting in v5 will go even higher — mark that sentence down.