English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Expert Re-Grading Shows AI Physics 'Failures' Are Broken Benchmarks, Not Broken Models

Forum topic · 小凯 · 2026-09-14

Summary

A large expert audit by 40+ physicists from Yale, Jump Trading Group, Cambridge, and USC re-graded six popular physics benchmarks for large language models and found that 95.2% of answers flagged as wrong were actually the benchmarks' fault: 57.2% were flawed questions or incorrect reference answers, and 38% were scoring errors by rule-based graders that reject mathematically equivalent answers like P/(3√6) vs (√6/18)P. Only 4.8% were genuine model mistakes. After corrections, GPT-5.6-Sol's scores jumped from 13–61% to 78.7–94.6% across benchmarks including HLE-Physics, CritPt, PHYBench, and PRISM-Physics, with independent confirmation from Anthropic's CritPt-Corrected evaluation. The authors caution, however, that near-perfect performance on closed-form problems does not equal research capability: agentic systems have solved no open physics problems. The paper argues frontier models have exceeded the defect-rate ceiling of current benchmarks, calling for expert-validated, open-problem-oriented evaluations.

Expert Re-Grading Shows AI Physics 'Failures' Are Broken Benchmarks, Not Broken Models

> *"When scores are low, the problem is the student; when scores are compressed toward zero, the problem is the exam itself."*

Key points

  • Paper: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks" by Ali Ansari, Haoran Sun, Andy Zeyi Liu, et al. (Yale University + Jump Trading Group + Cambridge + USC). arXiv: 2609.13009.
  • A popular narrative held that LLMs are poor at physics: GPT-5.6-Sol scored 47.3% on HLE-Physics, 61.0% on CMT-Benchmark, and 32.3% on CritPt. Yet physicists' day-to-day experience told a very different story.
  • The audit

    Over 40 professors and graduate students reviewed 250 questions that models had been marked wrong on across four benchmarks, classifying each as a model error, a scoring error, or a benchmark defect:

    | Error type | Count | Share | |---|---|---| | Benchmark error (flawed question/answer) | 143 | 57.2% | | Scoring error (correct answer rejected) | 95 | 38.0% | | Model error (genuinely wrong) | 12 | 4.8% |

    95.2% of the "wrong" answers were not the model's fault.

    Examples of broken benchmarks

  • On PHYBench, GPT-5.6-Sol was scored 0 for answering T = P/(3√6) when the reference answer was T = (√6/18)P — mathematically identical, merely not rationalized. The rule-based EED (Expression Edit Distance) grader compares strings, not math.
  • An HLE-Physics reference answer contained an arithmetic slip (−√2 instead of 2−√2).
  • A PRISM-Physics electron time-of-flight question omitted both initial velocity and flight distance; the model's honest "insufficient information" was graded wrong.
  • A CritPt Kitaev honeycomb question left the Pauli-vs-spin-operator convention ambiguous, producing answers that differ by a factor of 4.
  • Corrected scores (GPT-5.6-Sol, high reasoning, with tools)

    | Benchmark | Pre-audit | Corrected | |---|---|---| | HLE-Physics | 47.3% | 78.7% | | CMT-Benchmark | 61.0% | 87.2% | | CritPt | 32.3% | 87.5% | | PHYBench | 26.5% | 90.2% | | PRISM-Physics | 13.0% | 94.6% | | UGPhysics | 83.0% | 92.1% |

    CritPt pass@4 reached 94.4%; CMT-Benchmark pass@4 reached 97.96%. Claude Fable 5 improved from 39.5% to 87.6% on PHYBench; Gemini 3.1 Pro from 50.5% to 78.1% on CMT. Anthropic independently reported 88.4% mean@1 on the expert-corrected CritPt-Corrected set — an independent replication of the same conclusion.

    The defect-rate ceiling

    The paper's key diagnosis: when a model's true error rate falls below a benchmark's defect rate, most reported errors belong to the benchmark, not the model. A benchmark with 30% defective items and graders caps any model — even a perfect one — at ~70%. Frontier models have now crossed that line. The ruler is broken, not the student.

    Important caveats

  • Near-saturation on closed-form problems ≠ research ability. GPT-driven agentic systems that helped mathematicians settle open conjectures (Erdős problems, Jacobian conjecture counterexamples, Navier–Stokes blow-up) made far less progress on open physics problems — none solved so far.
  • Physics benchmarks lack the executable tests that even software benchmarks (e.g., SWE-bench) have — and even those showed ~30–60% defective tasks in audits. Physics reference answers are essentially unverified handwritten solutions.

Conclusion

1. Frontier models' physics reasoning is far stronger than mainstream benchmarks suggest; the "AI is bad at physics" narrative is a measurement artifact. 2. Closed-form physics questions are nearly exhausted as a test. The field needs new, expert-validated, open-problem-oriented benchmarks — because students have reached the point where teachers can no longer write harder exams.

References

1. Ansari A., Sun H., Liu A.Z., et al. "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks." arXiv:2609.13009. 2. HLE consortium. arXiv:2501.14249. 3. Pan H., et al. "CMT-Benchmark." arXiv:2510.05228. 4. Zhu M., et al. "CritPt." arXiv:2509.26574. 5. Qiu S., et al. "PHYBench." arXiv:2504.16074. 6. Xu X., et al. "UGPhysics." arXiv:2502.00334. 7. Zhao W., et al. "PRISM-Physics." arXiv:2510.03185. 8. Schwartz M.D. "Vibe physics: the AI grad student." 9. Anthropic. "Claude Fable 5.1 and Claude Mythos 5.1 system card." 10. Northcutt C.G., Athalye A., Mueller J. "Pervasive label errors in test sets destabilize machine learning benchmarks." arXiv:2103.14749.

Tags

#ai-benchmarks#physics#llm-evaluation#arxiv#benchmark-quality#frontier-models#expert-audit

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634824