A comprehensive survey of AI-assisted scientific research—reviewing 250+ papers from institutions including the National University of Singapore, the Chinese University of Hong Kong, and CNRS—diagnoses the state of the field as of April 2026. Its core conclusion: AI has become very good at producing the *forms* of research, but remains poor at guaranteeing the *substance*.
Key points
- Production is cheap and fast: The AI Scientist (2024) generates end-to-end papers for ~$15 each; FARS (2025) ran continuously for 228 hours, consuming 11.4 billion tokens to produce 100 papers (2.3 hours per paper); ARIS (2026) runs 20+ GPU experiments overnight and improved draft scores from 5.0 to 7.5.
- A unified framework: The survey models research as 4 phases and 8 steps — Idea Generation, Literature Review, Coding & Experiments, Tables & Figures, Paper Writing, Peer Review, Rebuttal & Revision, and Paper2X dissemination.
- Ideas look novel but aren't: LLM-generated ideas score high on novelty (>0.6 on IdeaBench) but low on feasibility (<0.5). HindSight found LLM-assessed novelty is *negatively* correlated with actual impact (ρ = -0.29).
- Research code is a weak spot: LLMs succeed at ~80% on pattern-matching tasks (SWE-Bench) but only 37-39% on research-level code (ResearchCodeBench) — a 41-43 point gap.
- Review failure is the most striking number: LLM reviewers misjudge 95.8% of papers that should be rejected as "acceptable," exhibit leniency bias, and can be manipulated by benign adjectives. 15.8% of ICLR reviews showed detected AI assistance, boosting borderline papers by +4.9pp.
- Writing is strong but flawed: CycleResearcher-generated papers scored 5.36 (vs 5.24 for preprints, 5.69 for accepted papers at ICLR 2025); citation accuracy (e.g., ScholarCopilot top-1 at 40.1%) remains far below human level.
- Dissemination (Paper2X) — posters, slides, videos, social media, interactive paper agents — is a maturing area, but fidelity issues persist.
- Use AI as a brainstorming partner, not an idea source; always verify novelty manually.
- Keep research-grade algorithm implementation, experimental design, and statistical choices under human control.
- Treat AI peer review as preliminary screening only; beware leniency bias.
- Core arguments and claims in papers should be human-written and validated.
- Survey: https://arxiv.org/abs/2605.18661
- Project page: https://worldbench.github.io/awesome-ai-auto-research
- The AI Scientist: https://www.nature.com/articles/s41586-024-07958-1
- OpenScholar: https://github.com/OpenScholar
- ResearchTown: https://github.com/ulab-uiuc/research-town
Five cross-cutting findings
1. AI is strong at structured, retrieval-driven tasks; weak at open-ended judgment, novel ideas, and experimental design. 2. Generation outpaces verification: a paper costs $15 and hours to make; validating reproducibility takes days to weeks. 3. Human-governed collaboration — not full autonomy — is the most reliable deployment mode. AI handles retrieval, drafting, and tooling; humans retain judgment, interpretation, design, and accountability. 4. Effective systems converge on layered architectures: exploration → tool execution → verification. Orchestration and provenance matter as much as model scale. 5. AI use is now a governance problem (disclosure, attribution, responsibility, integrity), not a detection problem.
End-to-end systems and open challenges
The survey categorizes systems into sequential pipelines (The AI Scientist, FARS), search & self-improving agents (Dolphin), skill/tool-integrated systems (MCP-based), and multi-agent community simulators (ResearchTown). None has consistently reached top-conference acceptance standards.
Eight open challenges are identified: cross-phase faithfulness, scientific judgment and novelty assessment, verification/reproducibility/accountability, citation and provenance, governance and research integrity, cross-domain generalization, human cognitive ownership, and building reliably trustworthy AI research assistants.
Practical recommendations
The ultimate takeaway: the core challenge is no longer whether AI can produce the forms of research, but whether it can preserve the substance — evidence, judgment, provenance, and accountability. Today, it cannot, which is why human-led human-AI collaboration remains the field's most reliable paradigm.