The Prophet's Dilemma: AI Tries to Predict the Future of Science and Falls Short — A Deep Dive into CUSP
| Paper Info | | |---|---| | Title | Forecasting Scientific Progress with Artificial Intelligence | | Authors | Sean Wu, Pan Lu, Yupeng Chen, Jonathan Bragg, Yutaro Yamada, Peter Clark, David Clifton, Philip Torr, James Zou, Junchi Yu (10 authors) | | Institutions | Shanghai Jiao Tong University (SJTU), Oxford, Stanford, Allen Institute for AI (AI2) | | arXiv ID | 2605.22681 | | Date | May 21, 2026 | | Category | cs.AI | | Scale | 4,760 scientific events; 73 pages; 13 figures; 29 tables | | Core claim | Current AI systems are far from reliable scientific oracles — they can identify promising directions but cannot reliably judge which will succeed or when, showing systematic overconfidence and strong response biases |
Edison once said, in essence: I haven't failed — I've found ten thousand ways that don't work.
The less romantic subtext: nobody knows which path will work until it has been walked.
Science is the same. Humanity invests roughly two trillion dollars in R&D annually. Most of that money goes down dead ends. It's not that scientists are incompetent — it's that the future itself is invisible. Stand in 2026 and ask someone: will AI surpass AlphaFold3 in protein structure prediction within two years? Will humans grow potatoes on Mars? Will room-temperature superconductors appear within five years?
They can only guess.
This paper asks precisely that question: Can AI guess more accurately than humans? The answer is surprising. It had ten frontier AI models across ten domains make predictions about 4,760 real scientific events. The result? The models are well-read, supremely confident — and firmly wrong.
The protagonist of the paper is a benchmark called CUSP — Cutoff-conditioned Unseen Scientific Progress. Three keywords in the name: Cutoff-conditioned (controlling the knowledge cutoff), Unseen, Scientific Progress. Together, the question it tests is brutally simple:
> AI, you've read all the papers up to 2021. Now tell me: between 2022 and 2025, which breakthroughs in your field will happen, when, and why.
The AI's answers resemble a student who copied half an answer sheet — but hasn't seen a single question on the actual exam.
---
🧪 How CUSP Was Built: A Library Assembled Backwards
To understand the study, you must first understand how CUSP was constructed.
4,760 scientific events across four major disciplines: AI, biology, chemistry, physics. Each event is a genuine scientific advance: the publication of a highly cited paper, a benchmark being broken, a technology moving from lab to industry.
But the key is not quantity — it's control. Each event carries a precise timestamp, and a knowledge cutoff is set. The background material given to models contains only literature published before the cutoff. Anything after that, the models have never seen and cannot retrieve.
The paper calls this a "temporally grounded evaluation framework." In plain terms: lock the AI inside a library of pre-cutoff knowledge, then ask it about tomorrow.
The library is meticulously built around four evaluation dimensions:
Feasibility Assessment: identify which candidate research direction is most likely to yield a breakthrough.
Mechanistic Reasoning: explain the underlying scientific principles and technical path behind a given advance.
Generative Solution Design: given an open problem, design a feasible research plan from scratch.
Temporal Prediction: estimate when a specific breakthrough will occur.
The first three test "understanding"; the fourth tests "prophecy." This four-dimensional design decomposes scientific forecasting into testable components — instead of vaguely asking "can AI predict science," it asks: at which stage does it break down?
---
📊 Four Findings, Each Sharper Than the Last
Four sets of results in this 73-page paper read like cold water to the face.
---
📉 Finding 1: Good at Ranking Directions, Terrible at Judging Outcomes
AI's performance on feasibility assessment is passable. Given a set of candidate directions, it can pick the most plausible ones. Having read enough papers, models know what a "serious-sounding" research question looks like.
But when asked "will this direction actually pan out" — it collapses.
The paper states that models fail to reliably predict whether scientific advances will be realized. This collapse is not random noise — it's systematic. Across domains and across models, the pattern repeats.
Judging direction: fine. Judging success or failure: no.
The distinction is chilling. Direction judgment relies on pattern matching — correct paper format, intact citation chains, clear problem statements, and the model recognizes it. Judging success requires modeling the physical world, experimental constraints, resource limitations, even the soft factor of "luck" — none of which are in the training data.
---
⏰ Finding 2: Time Is the Biggest Blind Spot
If "will it work" is hard, "when will it work" is absurdly harder.
The paper finds that models systematically misestimate when advances will occur. We're not talking about being off by months — the error can be off by an order of magnitude. In some domains, some models' temporal predictions look like dice rolls.
Behind this lies a deep dilemma. Breakthrough timelines are governed by countless uncontrollable factors: equipment breaks and takes a month to repair, a key collaborator leaves, a pandemic disrupts plans, reviewers take eight months — randomness that appears in no paper, no review, and that models can never learn.
You can read ten thousand papers and still not know when that mass spectrometer will break down.
---
🏭 Finding 3: AI Predicting AI — An Expert in the Mirror
Of the four domains, models perform best at predicting AI's own progress.
This is nearly tautological. AI models are trained on AI papers, AI code, AI paradigms — they have finer internal representations of their own field's timeline and milestones. AI progress is also more "industrialized": compute doubling times are predictable, benchmark iteration rhythms are trackable, open-source community activity is logged in real time.
Biology, chemistry, physics are different. Experimental cycles are slow, data generation takes longer, discoveries are often isolated results from a few labs, and externally visible signals are far sparser. Model forecasting accuracy on these "slow sciences" drops sharply.
"As the amount of pre-cutoff knowledge increases, model performance improves, but the gap to the 'full-information' setting never closes."
The subtlety here: showing the model every paper before the cutoff doesn't help — because the truly decisive information isn't in published papers. It's in a PI's unfinished manuscript, an unanswered collaboration email, a cryo-EM that's just been switched on.
---
🔍 Finding 4: Knowing Is Not the Same as Having Seen
The paper runs an elegant comparison: the same events, predicted under two conditions:
- Pre-cutoff knowledge only: literature before the cutoff.
- Full information: relevant post-cutoff information (essentially "hindsight mode").
Translated: what the models essentially do is not "predicting" but "recapping." When they can see "what happened later," they can narrate causes and effects fluently, convincingly. But cover up "what happened later," and they fall apart.
The paper's own words: "performance benefits more from post-event information than from forward-looking prediction."
And this pattern — seeming reasoning, actually retrieval — runs through all four evaluations.
---
🎭 Finding 5: Overconfident, Unaware of What It Doesn't Know
On top of all the poor performance is a more troubling layer: the models don't know they're guessing.
The paper reports "systematic overconfidence and strong response biases" — when models give wrong answers, they often attach extremely high confidence. It's hallucinating, but it believes it's right.
Confident errors are far more dangerous than cautious ones. You discount the judgment of a forecaster who knows it's uncertain. A forecaster who is wrong yet pounds the table will lead you into a ditch.
In scientific forecasting, this overconfidence is especially lethal. Human scientists are already naturally over-optimistic — our habit of assuming a technology is "five years away," combined with model overconfidence, produces dangerous resonance.
---
🔬 What's Trustworthy: Why This Research Is Solid
The paper's reliability lies not in startling conclusions but in the quality of the work.
First, massive sample size. 4,760 events, not a casually chosen few dozen. At this scale, statistical conclusions are reproducible.
Second, complete domain coverage. AI, biology, chemistry, physics — the core natural sciences. Not just benchmark-chasing on machine learning.
Third, rigorous temporal control. Cutoff-conditioning precisely controls the models' knowledge boundaries. This addresses a common flaw in AI evaluation: many seemingly "predictive" models actually rely on training data leakage.
Fourth, multi-dimensional, multi-model. Not one capability, one model — it's four dimensions × ten models. The conclusions are cross-validated, not propped up by a single metric.
The paper's stance is honest: not "we prove AI cannot forecast science," but "we measured the real upper limit of current AI systems on this task and localized the failure modes."
The distance between those two sentences is the distance between a good paper and clickbait.
---
❓ Honest Uncertainty: Questions This Paper Doesn't Answer
Some things I still don't know after reading.
Prediction or recall? The paper shows AI performs poorly when future information is hidden. But training data leakage is nearly impossible to fully exclude in such studies. Even if a model never directly "saw" a 2023 breakthrough, it may have indirectly, partially "foreknown" it — in a 2021 review's outlook, in conference slides, in a Twitter thread. The paper cannot control for this "soft leakage."
Where is the human baseline? No direct comparison with human experts. How bad are human scientists at forecasting? Perhaps under the same conditions, human experts aren't much better. Without a human baseline, the weight of "AI fails" is diminished.
Does CUSP's event sample have selection bias? How were the 4,760 events selected? If they are "breakthroughs that already succeeded," the test set carries positive selection bias — excluding the ten thousand paths that never worked. Testing "can AI recognize breakthroughs that already happened" versus "can AI predict a breakthrough before it happens" are entirely different questions. The cutoff-conditioned design attempts to address this, but the event set's construction deserves scrutiny.
What explains domain imbalance? Is AI-domain forecasting more accurate because AI advances are "engineering progress" rather than "scientific discovery"? Or because AI information is more digital, more public, more easily captured? The paper doesn't elaborate — and the cause determines how far the conclusion generalizes.
What causes the overconfidence? Is it an intrinsic bias of the training objective (maximizing next-token probability), or reward hacking from RLHF? The paper identifies the phenomenon but doesn't trace the root. Fixing it may require changing the training paradigm itself.
---
🦾 Stepping Back: What This Paper Really Weighs
After finishing, I sat and thought for a long time.
This is not a "AI fails again" rant. Its weight points the other way.
Scientific forecasting is not an ordinary AI capability. It is a meta-capability among meta-capabilities. If a system could reliably forecast scientific progress — which direction breaks through, which method succeeds, how long results take — it would have grasped the core laws of human knowledge production.
The paper shows today's AI is invisibly far from that goal.
But this may not be AI's fault. Forecasting science may simply be extremely hard — perhaps impossible.
Scientific progress depends on three kinds of information. First, published knowledge — papers, patents, public datasets. Models can read these. Second, unpublished but accessible knowledge — lab notebooks, internal discussions, failure records. Models cannot read these. Third, random events that haven't happened yet — instrument failures, personnel moves, policy changes, flashes of inspiration. No system can read these, because they don't yet exist.
If a substantial share of scientific progress's variance comes from the third kind — the paper doesn't discuss this, but I suspect the answer is "yes" — then the ceiling of scientific forecasting is far lower than anyone thinks.
Humans are bad at this too. Corporate labs launch hundreds of projects annually; single digits reach commercialization. Pharma spends an average of $2.6 billion per approved drug, with over 90% of candidates dying in preclinical or clinical stages. Grant reviewers' scores — their "predictions" — have repeatedly been shown to be barely better than random.
If humans and AI are both bad at forecasting scientific progress, what exactly are we comparing against?
Perhaps the paper's real message is not "AI fails," but a different question: have we long overestimated science's predictability?
---
💭 Coda: From "Prediction" to "Exploration"
One thing in the paper held my attention.
Not in the methods or the experiments — in the word choice. Across all 73 pages, the paper repeatedly uses "forecast" rather than "predict." The difference between these English words is small, but the direction is sharp.
Predict is "I compute" — a closed system where given inputs yield outputs. Weather forecasting is prediction — tomorrow's pressure field can be computed.
Forecast is "I estimate" — the system isn't fully closed, but based on the best available information and judgment, you give an estimate with uncertainty. Economic forecasting is forecasting — no one can compute next year's GDP, but you can give a range and a confidence level.
The paper's choice of forecast is not casual. It is a methodological acknowledgment of the boundary of prophecy: the scientific future cannot be precisely computed. But it might be reasonably estimated — if you truly understand the sources and magnitudes of uncertainty.
This leads to a deeper question: if science can't be predicted, what should AI's role in science be?
Not prophecy. Exploration. Not telling humans "where this road leads," but walking the paths humans can't — reading contradictions across ten thousand papers, finding patterns invisible to human intuition, exhausting possibilities in massively parallel experiments, compressing decades of trial and error into days of search.
The paper doesn't directly discuss this future. But it provides the premise for the discussion — first honestly measure the ceiling, then talk about what comes next.
That is genuine scientific spirit. And it is the most memorable thing about this 73-page paper.
---
📚 References
1. Wu, S., Lu, P., Chen, Y., Bragg, J., Yamada, Y., Clark, P., Clifton, D., Torr, P., Zou, J., & Yu, J. (2026). Forecasting Scientific Progress with Artificial Intelligence. *arXiv:2605.22681*. 2. Lu, P., Qiu, L., Chang, K. W., et al. (2024). ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. *arXiv:2410.05080*. 3. Chowdhery, A., Narang, S., Devlin, J., et al. (2023). PaLM: Scaling Language Modeling with Pathways. *JMLR*, 24(240):1-113. 4. Grace, K., Salvatier, J., Dafoe, A., Zhang, B., & Evans, O. (2018). When Will AI Exceed Human Performance? Evidence from AI Experts. *Journal of Artificial Intelligence Research*, 62, 729-754. 5. Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. *Crown Publishers*.