Overview
This post from zhichai.net, titled "The Invisible Hand: When AI Becomes a Co-Author in Science," reviews a large-scale randomized controlled trial on LLM feedback in scientific publishing.
Paper details:
| Item | Detail | |------|--------| | Title | Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment | | Authors | Binglu Wang, Weixin Liang, Jiahui Xue, Yuhui Zhang, Hancheng Cao, Dashun Wang, Yian Yin | | arXiv ID | 2605.24180 | | Submitted | May 22, 2026 | | Categories | physics.soc-ph; cs.AI; cs.DL; cs.HC |
Key Points
- Scale of the experiment: Over 31,000 arXiv preprints across 150 disciplines, involving 45,000+ researchers from 133 countries, were randomly assigned to receive either a structured, LLM-generated feedback report tailored to the paper, or nothing.
- Headline result: The treatment group's manuscript revision rate was 12.55% higher (relative increase) than the control group's baseline. The author argues this is a causal, reproducible behavioral shift driven purely by text feedback from a machine with no prior relationship to the authors.
- Who benefited most: Effects were strongest among researchers in non-English-dominant regions, papers with low citation-network embeddedness, teams with lower average h-index, and early-career researchers (PhD students, postdocs) — precisely those with the least access to quality human feedback.
- Spillover effect: Authors who received AI feedback showed a measurable increase in LLM tool usage in their subsequent publications, suggesting a second-order effect on scientific practice.
- Conceptual contribution: The paper reframes scientific feedback — historically a private good flowing through mentorship and social networks — as a distributable resource that does not require social capital to access.
- Careful claim scoping: The experiment compared AI feedback to no feedback, not AI feedback to human feedback; it does not claim AI feedback is superior to human review.
- The absolute baseline revision rate is unreported, making the practical significance of the 12.55% relative increase hard to assess.
- Revision quality was not measured — revision does not necessarily mean improvement.
- The false-positive rate of LLM feedback (incorrect suggestions) was not analyzed.
- Long-term effects on LLM adoption habits and disciplinary heterogeneity across the 150 fields remain open questions.
Honest Limitations Highlighted by the Author
Contextual Discussion
The author connects this work to a related paper (arXiv:2605.26340, ScientistOne) which found systematic fabrication (21% fake citations, 42% irreproducible benchmark scores) when LLMs autonomously execute the full research pipeline. The contrast suggests an emerging principle: the role AI plays in science determines its effect — fully autonomous AI may degrade scientific quality, while AI as an assistant (providing feedback without replacing human thinking) can improve equity and redistribute scientific resources.
References
1. Wang, Liang, Xue, Zhang, Cao, Wang & Yin, "Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment", arXiv:2605.24180, 2026. 2. Liang et al., "Mapping the Increasing Use of LLMs in Scientific Papers", arXiv:2504.02471v1, 2025. 3. Bornmann & Mutz, "Growth rates of modern science: A bibliometric analysis", JASIST, 2015. 4. Evans & Foster, "Metaknowledge", Science, 2011. 5. Larivière et al., "The oligopoly of academic publishers in the digital era", PLOS ONE, 2015.