Summary
Verdict: No indicators of academic fraud detected. The reviewed paper, published in Nature (Vol. 648) on 22 October 2025 by Oh, Farquhar, Kemaev, Calian, Hessel, Zintgraf, Singh, van Hasselt and Silver (Google DeepMind), describes the discovery of state-of-the-art reinforcement learning algorithms via meta-learning. Four review axes were applied because wet-lab checks (e.g. Western blot manipulation) are not applicable to this computational work. First, compute and timeline claims (1,024 TPUv3 cores for 64h for Disco57; 2,048 TPUv3 cores for 60h for Disco103) are realistic given DeepMind's infrastructure and the submission-to-acceptance window. Second, statistical reporting uses IQM with 95% confidence intervals across multiple random seeds (Atari/ProcGen/DMLab = 2; Crafter/NetHack = 3; Sokoban = 5), without evidence of cherry-picking. Third, the ablation study in Figure 3c yields a logically monotonic performance decline, consistent with the role of each component. Fourth, full code and meta-parameters are released under an open-source license, with conflict-of-interest disclosure. Confidence is moderate-to-high; final determination remains with institutional investigation.
Verdict
No specific fraud indicators detected (✅ Clean). The paper is computational (no wet-lab imagery), and the four review dimensions applied (compute/timeline plausibility, statistical reporting, ablation logic, open-source transparency) all returned normal.
Key findings
- Compute & timeline plausibility: Compute footprints (1,024 TPUv3 cores × 64h for Disco57; 2,048 TPUv3 cores × 60h for Disco103) are consistent with Google DeepMind's cluster scale over the 2024-12-11 submission to 2025-10-22 acceptance window.
- Statistical rigour: The work uses IQM (Interquartile Mean) with 95% confidence intervals as recommended by Agarwal et al., 2021 (reference 29). Random seed counts are explicitly disclosed: Atari/ProcGen/DMLab = 2 seeds, Crafter/NetHack = 3 seeds, Sokoban = 5 seeds. No indication of cherry-picking.
- Ablation logic (Figure 3c): Performance decreases monotonically with component removal: full Disco57 = 11.6; remove auxiliary prediction = 10.5; smaller agent network = 8.7; remove newly discovered y/z prediction = 5.0; meta-learn in simple gridworld = 3.4; remove value function q = 2.9. The steepest drop after removing the value function is consistent with RL theory, suggesting internally coherent data rather than fabricated trends.
- Open-source & disclosure: Code and Disco103 meta-parameters are openly released (https://github.com/google-deepmind/disco_rl). Conflicts of interest (Google LLC commercial interests and pending patents) are disclosed in the manuscript.
Evidence highlights
- DOI: 10.1038/s41586-025-09761-x
- Performance values cited from Figure 3c (Disco57 ablation): 11.6, 10.5, 8.7, 5.0, 3.4, 2.9
- Compute figures from Methods – Implementation details: 1,024 TPUv3 × 64h; 2,048 TPUv3 × 60h
- Seed counts: Atari/ProcGen/DMLab = 2; Crafter/NetHack = 3; Sokoban = 5
- Statistical method: IQM with 95% confidence intervals, citing Agarwal et al., 2021
- Code release: github.com/google-deepmind/disco_rl
Notes
- Wet-lab image checks (Western blot splicing, gel manipulation) are not applicable to this computational paper and were intentionally skipped.
- Confidence level: moderate-to-high. The transparency measures (open-source code, disclosed seeds, IQM with CIs) materially reduce, but do not eliminate, the possibility of undisclosed issues.
- Limitations of this automated review: we did not execute the released code to independently reproduce reported scores; potential issues such as hyperparameter selection bias or unreported training variance cannot be ruled out without full reproduction.
- Final determination of misconduct requires institutional or editorial investigation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a2f09f02bf156.93620411