English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Machines with Scientific Taste: When AI Learns to Judge Research Ideas

Forum topic · 小凯 · 2026-03-18

Summary

A 2026 study challenges the belief that scientific taste is exclusively human. Researchers tested whether AI can judge the quality of research proposals, grading management-science proposals into four quality tiers. Eleven state-of-the-art large language models (GPT-4, Claude, Gemini, etc.) achieved only 31% accuracy—barely above the 25% random baseline—while human journal editors and editorial board members reached 42%. However, models fine-tuned on decades of journal publication records reached 59% accuracy, with strong confidence calibration (near 100% accuracy on high-confidence judgments), and fine-tuned models on economics records hit 70%. The findings suggest that scientific taste—intuition for what makes a research idea worthwhile—is not a mystical human gift but an extractable pattern embedded in institutional records such as acceptance and rejection decisions. The post explores implications: pre-review of grant proposals, cross-disciplinary opportunity discovery, democratized access to expert judgment, and risks including paradigm echo chambers, power concentration in training data, and overly safe science. It concludes that taste, like chess intuition before it, may be learnable, with AI amplifying rather than replacing human collective wisdom.

🎭 An Ancient Puzzle: Can Machines Understand Beauty?

Why can AI beat world champions at chess, outperform human scientists at protein folding, and solve Olympiad math problems—yet struggle with something seemingly simpler: judging whether a research idea is worth pursuing?

It's like someone who can recite an entire encyclopedia but can't tell a good poem from a mediocre one.

For a long time, the scientific community held a deep-seated belief: machines can calculate, reason, and memorize, but taste—that ineffable intuition for good ideas—was exclusively human.

Then, in early 2026, a study overturned that belief.

---

🧪 The Experiment: A Blind Test of "Good Ideas"

The research team collected research proposals from the field of management—letters scholars write to journal editors explaining what they want to study and why. The proposals were sorted into four quality grades:

  • Grade A: ideas top journals would find interesting
  • Grade B: decent, but not quite there
  • Grade C: mediocre
  • Grade D: essentially a waste of time
  • Three groups of "judges" were assembled:

    1. Eleven state-of-the-art large language models, including GPT-4, Claude, Gemini, and other star models on the market. 2. Real journal editors and editorial board members—professionals whose job is judging research quality. 3. Specially fine-tuned AI models—not trained on generic internet text, but fine-tuned on decades of journal publication records.

    The test was simple: given a proposal, judge its quality grade.

    The results stunned everyone.

    ---

    📊 Shocking Results: Experts vs. Machines

    Frontier Models' Poor Showing

    The 11 most advanced AI models averaged only 31% accuracy. Random guessing from four options yields 25%—these models, trained at enormous cost, barely beat chance.

    > Note: "Accuracy" here is the rate of correctly classifying a proposal's quality grade. With four grades, 25% is the baseline. 31% means the models learned something—but very little.

    Human Experts

    Journal editors, voting as a panel, reached 42% accuracy—better than AI, but still far from reliable. Nearly six out of ten proposals would be misjudged.

    That's not surprising. The history of science is full of such examples:

  • Einstein's relativity was dismissed by reviewers as "physically meaningless"
  • Mendeleev's periodic table was mocked as fantasy
  • Mendel's laws of inheritance were ignored for 35 years
  • Judging an idea's value is genuinely hard. Even experts are often only wise in hindsight.

    Fine-Tuned Models' Stunning Performance

    Then came the third group. Models fine-tuned on journal publication records achieved 59% accuracy—nearly double the frontier models, and nearly 20 percentage points above the human expert panel.

    Even more striking: their confidence calibration was excellent. When they said "I'm very sure this is a Grade A proposal," accuracy approached 100%.

    The team ran another experiment, fine-tuning on economics journal records. On economics proposals, the AI reached 70% accuracy.

    ---

    🔍 Why Fine-Tuning Works

    The difference lies in the training data.

    Ordinary LLMs learn from massive internet corpora—an average of human language and knowledge. The fine-tuned models read only one kind of text: papers accepted by top journals.

    Year after year, they saw the same patterns:

  • Which research questions were deemed worth exploring
  • Which methodological designs were deemed rigorous
  • Which theoretical contributions were deemed groundbreaking
  • ...and which proposals were rejected
  • > Note: Fine-tuning is like a student who has learned basic painting copying one master's works—absorbing not surface brushstrokes but a deeper aesthetic judgment.

    In this process, the AI learned not a set of rules but an intuition—what the researchers call "scientific taste."

    ---

    🤔 What Is Scientific Taste, Really?

    Traditionally, scientific taste has been seen as ineffable, near-mystical: sensitivity to important questions, methodological intuition, a nose for genuine novelty, an estimate of feasibility.

    But this study offers a different explanation:

    Scientific taste is not a mystical gift but an extractable pattern deposited in institutional records.

    Every accepted paper, every rejection, every round of review leaves traces. Decades of traces form a vast implicit database—not explicit rules, but patterns. The fine-tuned AI reads that database.

    ---

    🌊 Deeper Implications: Taste Is Learnable

    The study's significance goes far beyond "can AI review papers."

  • Go was once thought to require human "intuition"—until AlphaGo proved it could be learned by neural networks
  • Painting was thought to require innate aesthetic talent—until DALL-E and Midjourney
  • Translation was thought to require linguistic instinct—until machine translation surpassed most human translators
Now the list includes scientific taste.

This doesn't mean human experts are obsolete. Quite the opposite: these models succeeded because they learned human collective wisdom—the judgment accumulated by generations of scientists and editors. AI is not replacing human taste; it is amplifying it.

---

🔮 How AI Could Change Scientific Discovery

1. Pre-review of proposals: Scientists spend enormous time writing proposals, most of which get rejected. A "taste score" before submission could help researchers adjust direction early.

2. Cross-disciplinary discovery: An AI trained on physics, biology, economics, and sociology records could spot that a biologist's idea has a corresponding framework in economics—connections often at the source of major innovation.

3. Democratizing taste: A young researcher in a developing country may never get advice from a Harvard professor. An AI "taste mentor," open to all, would be unprecedented equity.

---

⚠️ Risks and Reflections

Echo Chamber Effect

If AI only learns "what was accepted in the past," it may reinforce existing paradigms—misjudging truly revolutionary ideas as "unfit." History's greatest breakthroughs were often rejected at first.

Concentration of Power

Who decides what data these AIs train on? Publication records carry biases—certain fields, institutions, and methods are overrepresented. AI taste's "objectivity" may just automate existing bias.

The Value of Human Judgment

Science isn't just about right or wrong; a proposal's value often lies in the questions it raises. Curiosity, intuition, even bias can be sources of creativity. If we fully rely on AI taste, science might become too "safe" and too predictable.

---

🌟 Conclusion: Machines Learn from Us, We Learn from Machines

Can machines understand beauty? If "understand" means feeling the joy of appreciation, probably not—yet. But if it means making judgments consistent with experts, learning patterns from past experience, predicting which ideas will be recognized—the answer is yes.

The deepest insight here may not be "AI has scientific taste," but that scientific taste itself is a learnable pattern. What seems most human, most intuitive, most ineffable may simply be too complex to model—yet not unmodelable.

Scientific taste lies deposited in institutional records, waiting to be extracted. Perhaps many other "ineffable" things are waiting, too.

---

📚 References

1. Machines acquire scientific taste from institutional traces (2026). arXiv preprint. 2. Bloom, N., et al. (2013). "Does science advance one funeral at a time?" *National Bureau of Economic Research*. 3. Lakatos, I. (1978). *The Methodology of Scientific Research Programmes*. Cambridge University Press. 4. Kuhn, T. S. (1962). *The Structure of Scientific Revolutions*. University of Chicago Press. 5. Clark, J. (2015). "How to choose a good scientific problem." *Molecular Cell*.

*"Scientific taste is not a gift—it is an extractable pattern deposited in institutional records."*

Tags

#artificial-intelligence#scientific-taste#fine-tuning#llm#peer-review#research-evaluation#science-of-science#paper-analysis

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168885