English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Integrity Review: "Faster sorting algorithms discovered using deep reinforcement learning" (Nature, 2023) – DOI 10.1038/s41586-023-06004-9

Academic fraud report · Geng Detector

Summary

This report reviews the Google DeepMind paper published in Nature (Vol. 618, 8 June 2023) for signs of academic fraud. The verdict is CLEAN: no evidence of data falsification, image manipulation, statistical fabrication, or methodological inconsistency was identified. Because this is a computer science / reinforcement-learning paper rather than a wet-lab biology study, the usual image-duplication and blot-reuse checks do not directly apply. However, a structured review was performed covering (a) text/typography extraction artefacts in figure captions and pseudocode panels, (b) internal mathematical and statistical consistency of the benchmark latency data in Table 1, (c) timeline coherence between submission dates, LLVM open-source patches, and TPU hardware availability, and (d) logical consistency of the exhaustive-search claims in the Methods. All four axes passed. Confidence is high because the derived algorithms were released as open source and integrated into the LLVM standard library, where independent reproduction by the global developer community would have surfaced any fabrication. Limits: this review does not substitute a formal institutional investigation.

Verdict

Clean – no indicators of academic fraud identified. The paper passes integrity checks on data consistency, timeline plausibility, and methodological self-consistency. The reported algorithms are publicly open-sourced in LLVM, providing strong external reproducibility.

Key findings

  • No image fabrication risk: Being an algorithms/AI paper, there are no Western blots, microscopy images, or comparable figures that could be reused or spliced.
  • Text-extraction artefacts only (not fraud): Apparent "gibberish" strings such as Y R L G Y D U L D E O H B V R U W B L Q W O H Q J W K near Figure 1 and the pseudocode regions of Figure 3 are PDF font-encoding / double-byte extraction artefacts; in the rendered PDF these correspond to legitimate C++ and assembly pseudocode. Attributed to journal typesetting, not author misconduct.
  • Benchmark statistics are realistic: Table 1b reports asymmetric confidence intervals (e.g., VarSort3 latency: 236,498 with interval 235,898–236,887; lower spread 600, upper spread 389), consistent with real micro-benchmark noise across 100 machines. No "unnaturally perfect" values were observed.
  • Timeline is coherent: Received 25 July 2022, Accepted 23 March 2023; the upstream LLVM patches (reference 3, Gelmi M., LLVM.org 2022) align with the sorting-algorithm disclosures. TPU v3 / v4 hardware was available at Google at the claimed times.
  • Methodology self-consistent: The claim of a 3-day exhaustive enumeration of ~10^32 programs to prove no sorting routine shorter than 17 instructions exists for n=3 is internally consistent with combinatorial-explosion arguments; no contradiction between claimed search space and resource expenditure was detected.
  • Open-source release lowers fabrication incentive: The sorting routines were integrated into the LLVM C++ standard library, exposing them to continuous community review.
  • Evidence highlights

  • DOI: 10.1038/s41586-023-06004-9
  • Asymmetric latency CI example (Table 1b): 236,498 with bounds (235,898, 236,887); spread lower = 600, upper = 389.
  • Submission dates: Received 25 July 2022, Accepted 23 March 2023.
  • Reference 3: Gelmi, M. LLVM.org 2022 (LLVM patch integrating Sort3/4/5).
  • Reported search-space size for exhaustive proof: ~10^32 programs over ~3 days.
  • Notes

  • Standard biology-focused detection dimensions (Western-blot reuse, gel splicing, etc.) are not applicable to this theoretical CS/AI paper; their absence is not evidence of innocence or guilt and should be read as scope-appropriate.
  • The review was AI-assisted and is intended for academic discussion only; it does not replace an institutional investigation.
  • Confidence in the clean verdict is high given independent open-source reproduction, but residual uncertainty (e.g., undocumented hyperparameter tuning, unreported negative results) is inherent to any desk review of a large ML paper and cannot be fully excluded.

Tags

#academic-fraud-screen#computer-science#reinforcement-learning#deepmind#nature#sorting-algorithms#clean-verdict#open-source-verification

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a2ef09abbc3e7.85986809