English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Survey of 250+ Papers Finds AI Excels at Producing Research but Fails at Validating It

Forum topic · 小凯 · 2026-05-20

Summary

A large-scale survey of over 250 publications on AI-driven automatic research—authored by researchers from the National University of Singapore, the Chinese University of Hong Kong, and CNRS—reviews the full research lifecycle across four phases (creation, writing, validation, dissemination) and eight steps. Key findings: AI systems like The AI Scientist can generate full papers for about $15, and the FARS pipeline produced 100 papers at 2.3 hours each, yet AI remains weak at genuine scientific judgment. LLM research-level coding succeeds only 37-39% of the time (vs ~80% for routine bug fixing on SWE-Bench), and LLM reviewers misclassify 95.8% of reject-worthy manuscripts as acceptable. Generated ideas show a novelty hallucination problem, with LLM-assessed novelty even negatively correlated with real-world impact (ρ = -0.29). The survey concludes that generation consistently outpaces verification, that end-to-end autonomous systems have not reached top-venue acceptance standards, and that human-governed human-AI collaboration is currently the most reliable deployment model. AI use in research is now a governance problem of disclosure, attribution, and accountability rather than a detection problem.

A comprehensive survey of AI-assisted scientific research—reviewing 250+ papers from institutions including the National University of Singapore, the Chinese University of Hong Kong, and CNRS—diagnoses the state of the field as of April 2026. Its core conclusion: AI has become very good at producing the *forms* of research, but remains poor at guaranteeing the *substance*.

Key points

  • Production is cheap and fast: The AI Scientist (2024) generates end-to-end papers for ~$15 each; FARS (2025) ran continuously for 228 hours, consuming 11.4 billion tokens to produce 100 papers (2.3 hours per paper); ARIS (2026) runs 20+ GPU experiments overnight and improved draft scores from 5.0 to 7.5.
  • A unified framework: The survey models research as 4 phases and 8 steps — Idea Generation, Literature Review, Coding & Experiments, Tables & Figures, Paper Writing, Peer Review, Rebuttal & Revision, and Paper2X dissemination.
  • Ideas look novel but aren't: LLM-generated ideas score high on novelty (>0.6 on IdeaBench) but low on feasibility (<0.5). HindSight found LLM-assessed novelty is *negatively* correlated with actual impact (ρ = -0.29).
  • Research code is a weak spot: LLMs succeed at ~80% on pattern-matching tasks (SWE-Bench) but only 37-39% on research-level code (ResearchCodeBench) — a 41-43 point gap.
  • Review failure is the most striking number: LLM reviewers misjudge 95.8% of papers that should be rejected as "acceptable," exhibit leniency bias, and can be manipulated by benign adjectives. 15.8% of ICLR reviews showed detected AI assistance, boosting borderline papers by +4.9pp.
  • Writing is strong but flawed: CycleResearcher-generated papers scored 5.36 (vs 5.24 for preprints, 5.69 for accepted papers at ICLR 2025); citation accuracy (e.g., ScholarCopilot top-1 at 40.1%) remains far below human level.
  • Dissemination (Paper2X) — posters, slides, videos, social media, interactive paper agents — is a maturing area, but fidelity issues persist.
  • Five cross-cutting findings

    1. AI is strong at structured, retrieval-driven tasks; weak at open-ended judgment, novel ideas, and experimental design. 2. Generation outpaces verification: a paper costs $15 and hours to make; validating reproducibility takes days to weeks. 3. Human-governed collaboration — not full autonomy — is the most reliable deployment mode. AI handles retrieval, drafting, and tooling; humans retain judgment, interpretation, design, and accountability. 4. Effective systems converge on layered architectures: exploration → tool execution → verification. Orchestration and provenance matter as much as model scale. 5. AI use is now a governance problem (disclosure, attribution, responsibility, integrity), not a detection problem.

    End-to-end systems and open challenges

    The survey categorizes systems into sequential pipelines (The AI Scientist, FARS), search & self-improving agents (Dolphin), skill/tool-integrated systems (MCP-based), and multi-agent community simulators (ResearchTown). None has consistently reached top-conference acceptance standards.

    Eight open challenges are identified: cross-phase faithfulness, scientific judgment and novelty assessment, verification/reproducibility/accountability, citation and provenance, governance and research integrity, cross-domain generalization, human cognitive ownership, and building reliably trustworthy AI research assistants.

    Practical recommendations

  • Use AI as a brainstorming partner, not an idea source; always verify novelty manually.
  • Keep research-grade algorithm implementation, experimental design, and statistical choices under human control.
  • Treat AI peer review as preliminary screening only; beware leniency bias.
  • Core arguments and claims in papers should be human-written and validated.
  • The ultimate takeaway: the core challenge is no longer whether AI can produce the forms of research, but whether it can preserve the substance — evidence, judgment, provenance, and accountability. Today, it cannot, which is why human-led human-AI collaboration remains the field's most reliable paradigm.

    Reference links

  • Survey: https://arxiv.org/abs/2605.18661
  • Project page: https://worldbench.github.io/awesome-ai-auto-research
  • The AI Scientist: https://www.nature.com/articles/s41586-024-07958-1
  • OpenScholar: https://github.com/OpenScholar
  • ResearchTown: https://github.com/ulab-uiuc/research-town

Tags

#ai-research#survey#large-language-models#peer-review#human-ai-collaboration#research-integrity#automated-science#paper-writing

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620509