If you send a top student to the International Physics Olympiad (IPhO), they might win gold. But put them in a lab to independently tackle a frontier research problem, and they may not even find the door.
The AI world faces the same split between the exam hall and the battlefield. OpenAI's recently released scientific benchmark FrontierScience (2026) delivers a sobering report card for large language models: AI is near-divine at answering questions, but before real scientific discovery, it is still a toddler learning to speak.
1. Two Faces of Science: Olympiad vs. Research
OpenAI argues that judging whether an AI truly "understands science" requires two measures:
- Olympiad track: Extremely difficult problems at the international elite level in physics, chemistry, and biology. This tests the model's problem-solving fundamentals—formula derivation and logical rigor.
- Research track: The real battlefield. It requires the model to act like a PhD student: forming hypotheses from ambiguous clues, designing experiments, and analyzing data no one has ever seen.
- Exam master: On the Olympiad track it scored 77%, meaning most human experts no longer hold any problem-solving advantage over it.
- Research novice: Yet on the Research track, its score immediately dropped to 25%.
- Hypothesis generation: AI excels at running along established paths but struggles to think outside the box and propose counterintuitive approaches.
- Error tolerance: Science is full of failure. Accustomed to finding the "correct path," when experimental data contradicts logic, AI often hallucinates a forced explanation instead of questioning its underlying reasoning.
- Long-horizon planning: A scientific experiment can run for months. Although AI's context length keeps growing, it remains weak at deep, multi-step decision-making across long time scales.
2. The Painful Report Card: GPT-5.2's Ceiling
Top-tier GPT-5.2's results are thought-provoking:
A Feynman-style analogy: Imagine a student who has memorized an entire physics textbook and can instantly solve the most complex projectile equations. But drop them on a desert island and ask them to invent a fluorescent lamp with available resources—beyond reciting the book, they can do nothing. That is AI's current state: it has top-tier problem-solving intuition but severely lacks scientific creativity.
3. Why AI Can't Replace Scientists Yet
The paper points to three soft spots of AI in real research:
Editorial Take
The arrival of FrontierScience marks a new era in AI training: a leap from imitating human conversation to solving humanity's unsolved mysteries.
A research score of 25 may look pathetic, but this is precisely the most fascinating, most promising uncharted territory of intelligence. When AI crosses this 50-point gap—from a machine that answers questions to a colleague that thinks—the pace of human civilization's progress will be rewritten.
If AI one day scores 100% on scientific research, what do you think human scientists should do? Start the ultimate debate in the comments!
---
*Note: This article is based on OpenAI's scientific evaluation benchmark released in May 2026.*