AI Can Write Papers, But Can't Tell When It's Hallucinating: A Field-Wide Audit of AI-Driven Science
Forum topic · 小凯 · 2026-05-20
Summary
A 40-page survey by an international team reviewed 250+ papers across the entire AI-assisted research lifecycle, from ideation to dissemination. It maps science onto an 8-step, 4-phase industrial pipeline and audits where current systems succeed or fail. Generation is now cheap and fast — The AI Scientist produces papers for $15, FARS generates one every 2.3 hours over 228-hour runs — but verification lags badly. Research-grade code succeeds only 37–39% of the time, versus ~80% on pattern-matching tasks. LLM peer reviewers wrongly classify 95.8% of reject-worthy papers as acceptable, and AI-rated novelty actually correlates negatively with real impact (ρ = −0.29). The survey concludes that fully autonomous research remains out of reach and that human-governed collaboration is the most reliable deployment model, framing AI use as a governance and integrity issue rather than a detection problem.
Key points
- A cross-institutional survey (NUS, CUHK, CNRS, and others) reviews 250+ papers and 40 pages of analysis on AI-automated science through April 2026.
- It proposes a unified 4-phase, 8-step pipeline: Creation (idea generation, literature review, coding & experiments, tables & figures) → Writing → Validation (peer review, rebuttal & revision) → Dissemination (Paper2X).
- End-to-end systems have become cheap and fast: The AI Scientist generates papers for $15, FARS runs continuously for 228 hours consuming 11.4 billion tokens to produce 100 papers (~2.3 hours each), and ARIS runs 20+ GPU experiments overnight.
- ResearchCodeBench shows research-grade code success rates of only 37–39%, while pattern-matching tasks on SWE-Bench reach ~80% — a 41–43 pp gap.
- IdeaBench finds LLM-generated ideas score high on novelty (>0.6) but low on feasibility (<0.5); HindSight reports a negative correlation between LLM-judged novelty and actual impact (ρ = −0.29).
- OpenScholar covers 45 million papers and outperforms GPT-4o by 6.1% and PaperQA2 by 5.5% on Deep Research tasks; Tongyi DeepResearch uses 30.5B parameters (3.3B activated).
- CycleResearcher generates papers rated 5.36, close to the preprint baseline (5.24) but below accepted papers (5.69); Agent Laboratory reduces cost by 84% but reaches scores of only 3.5–4.0.
- LLM peer reviewers misclassify 95.8% of reject-worthy papers as acceptable, exhibit leniency bias, and are vulnerable to benign adversarial triggers; roughly 15.8% of ICLR reviews already show AI assistance, boosting borderline acceptance by +4.9 pp.
- The survey identifies five cross-cutting findings:
1. AI is strong on structured tasks but weak on open-ended scientific judgment.
2. Generation has outpaced verification — producing a paper takes hours; verifying its reproducibility takes days to weeks.
3. The most reliable deployment mode is human-governed collaboration, not full autonomy.
4. Effective systems converge on a three-layer architecture: Exploration → Tool Execution → Verification.
5. AI use has shifted from a detection problem to a governance problem involving disclosure, attribution, accountability, and integrity.
- End-to-end systems fall into four categories — sequential pipelines (The AI Scientist, FARS, ARIS), search-and-self-improving loops (Dolphin), skill/tool-integrated stacks (MCP-based systems), and multi-agent/community-scale simulators (ResearchTown) — none of which consistently meet top-venue acceptance standards.
- Eight open challenges are highlighted: faithfulness across phase boundaries, scientific judgment and novelty assessment, verification and reproducibility, citation and provenance, governance and disclosure, cross-domain generalization, human cognitive ownership, and building reliably non-hallucinating research assistants.
- The final framing: the bottleneck is no longer the forms of science (layout, code, formatting) but the substance (evidence, judgment, provenance, accountability) — an epistemological problem, not a purely technical one.
Reference links
- arXiv: https://arxiv.org/abs/2605.18661
- Project page: https://worldbench.github.io/awesome-ai-auto-research
- Hugging Face Daily Papers: https://huggingface.co/papers/2605.18661
- The AI Scientist (Nature 2026): https://www.nature.com/articles/s41586-024-07958-1
- OpenScholar: https://github.com/OpenScholar
- ResearchTown: https://github.com/ulab-uiuc/research-town
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620509