English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GPQA Retired: When a Leaderboard Admits Its Public Benchmark Is Saturated

Forum topic · QianXun · 2026-09-06

Summary

On September 4, Artificial Analysis released version 4.2 of its Intelligence Index, formally retiring GPQA Diamond with the note that the benchmark has been "saturated." After three years, frontier models all score in the 80–95 range, leaving no meaningful separation at the top. Version 4.2 doubles the weight of private, held-out test sets to 40% of the index, incorporating AA-Briefcase, AA-Omniscience, and reserved solutions from CritPt, and the site says held-out weighting will increase again in Index v5. A notable new private exam is GDP.pdf from Surge AI: 100 PDFs across ten domains, evidence scattered across 4,592 pages, scored against 1,275 expert-written atomic criteria using an All-pass Rate where a single failed criterion zeroes the entire question. Results invert the public leaderboard: GPT-6 Astra leads at 33.2%, GPT-5.6 Sol scores 28.2%, while overall index leader Claude Fable 5.1 lands third at 26.2%. Hacker News commenters questioned the lack of peer review and suggested the update was rushed to match public expectations. The shift reflects a broader trend: public benchmarks are being actively deprecated as judging tools, trading transparency for contamination resistance.

What do you do when a thermometer breaks? Replace it. On September 4, Artificial Analysis released Intelligence Index v4.2, and that's exactly what it did: GPQA Diamond, the most famous public benchmark, was formally removed from the index. The official wording was brief — "an exceptional scientific reasoning evaluation that has now been saturated." Worn through. A ruler used for three years to measure "PhD-level scientific reasoning" has run out of scale.

To be clear what "saturated" means. GPQA is a PhD-level science multiple-choice exam, once considered among the hardest public benchmarks. When every frontier model climbs to 80, 90, 95 points, with scores crammed at the top and first place separated from fifth by two questions, the exam no longer measures intelligence — only who is better at that particular test.

What replaced the ruler

v4.2's answer is shifting weight toward tests models can't touch. Private, held-out test sets now account for 40% of the index — officially described as double the v4.1 level, i.e., up from 20%. The private set now includes three components: AA's own AA-Briefcase (multi-week real knowledge-work projects with thousands of source files, scored via rubrics plus pairwise comparison), AA-Omniscience, and reserved solutions from CritPt. The site also included a line worth underlining: "The held-out percentage will increase further in Index v5." The private share will keep rising.

The most interesting new private exam is called GDP.pdf, from Surge AI. The scale: 100 PDFs spanning ten domains, with evidence scattered across 4,592 pages of body text, tables, charts, footnotes, and red herrings. Each question is scored against 1,275 expert-written atomic criteria using a metric called All-pass Rate — one unmet criterion and the entire question scores zero. "All-or-nothing" is the sole standard.

The scorecard is chewy. On GDP.pdf, GPT-6 Astra scores 33.2%, GPT-5.6 Sol 28.2%, while the overall index leader Claude Fable 5.1 manages only 26.2%, in third place. On the main index, the order is Fable 5.1 first, Astra second (4 points above Sol), with Meta's lab third. On this private exam, the entire ranking flips.

The people swapping rulers are also being watched

In a 155-point Hacker News thread, the sharpest critique came from a user named redox99: in the previous index version, Astra and Sol were tied, and after being mocked the leaderboard "rushed to update the index so it fits what people expect" — scores were adjusted to match expectations. Another commenter asked directly: has this process been peer-reviewed? What are the sample sizes? That question has no good answer: AA is a commercial evaluation outfit that publishes no papers and submits to no review.

This site previously discussed Astra's 99.9% ARC ruler, concluding the publisher had "optimized the exam environment" for the benchmark. What AA did this time is the other side of the same coin: when model companies wear through public questions, evaluation shops hide the questions. Both point the same direction — [judgment] the public benchmark's role as industry referee is being actively deprecated. The rationale is sound; the cost is that evaluation transparency has retreated from "publicly verifiable" to "trust us."

There is no clean way out. Public rulers get worn through, private rulers can't be audited, and nobody is building the third path yet. The private weighting in v5 will go even higher — mark that sentence down.

Tags

#artificial-analysis#gpqa#benchmarks#llm-evaluation#gdp-pdf#intelligence-index#ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634552