Summary
Verdict: No evidence of academic fraud detected. This computational biology/machine learning paper by Schulte-Sasse, Budach, Hnisz and Marsico, published in Nature Machine Intelligence, was reviewed across images, statistical reporting, data consistency, and citation timeline. No image duplication or splicing issues could be identified, though pixel-level verification was constrained by text-only extraction. Reported statistics (Spearman ρ = 0.63, P < 2.2 × 10⁻¹⁶, P = 1.6 × 10⁻¹⁵, P = 4.9 × 10⁻¹¹) show authentic computational characteristics, including the typical R statistical-software minimum p-value output rather than artificially rounded figures. The paper explicitly notes that of 165 newly predicted cancer genes, only 163 could be included in the contingency-table analysis due to missing entries in the CRISPR screen database—an authentic sign of real data handling. No p-hacking indicators were observed. All software, databases (GENCODE v.28, TCGA, GISTIC2, Project Achilles), and cited literature (up to February 2020) are temporally consistent with the July 2020 submission date. Confidence is moderate due to text-only access; final conclusions require official institutional investigation.
Verdict
No evidence of academic fraud detected. The paper appears methodologically sound, statistically authentic, and internally consistent.
Key findings
- Image analysis limited by text-only extraction; no logical contradictions or visual duplications identified in Figures 1–6.
- Statistical reporting shows authentic computational characteristics (Spearman ρ = 0.63, P < 2.2 × 10⁻¹⁶, P = 1.6 × 10⁻¹⁵, P = 4.9 × 10⁻¹⁹).
- Transparent reporting of data incompleteness: 165 predicted cancer genes reduced to 163 in contingency-table analysis due to missing CRISPR screen database entries.
- No p-hacking signals (no clustering of p-values near the 0.04–0.05 threshold).
- Rigorous cross-validation methodology (grid search with fivefold cross-validation; external test sets OncoKB, ONGene).
- Citation timeline fully self-consistent with July 2020 submission (latest cited references from February 2020).
- Methods cite established tools and databases (GENCODE v.28, TCGA, GISTIC2, Project Achilles) all available prior to the submission date.
Evidence highlights
- DOI: 10.1038/s42256-021-00325-y
- Submission: 21 July 2020; Accepted: 18 February 2021; Published online: 12 April 2021
- Reported Spearman correlation: 0.63 with P < 2.2 × 10⁻¹⁶, consistent with R's standard minimum p-value output.
- 165 newly predicted cancer genes (NPCGs); 163 retained for Figure 4c contingency-table analysis due to missing database entries.
- Open code and data sharing via GitHub repository (as reported by the reviewer).
Notes
- Limitations: This assessment relied on text extraction only; pixel-level image forensics, source data file auditing, and code execution checks were not performed.
- Image duplication detection for bioinformatics heatmaps, network diagrams, and model architecture schematics requires original high-resolution figure files for definitive conclusions.
- No additional action required based on available evidence. Any final determination of academic misconduct must be made by qualified institutional investigation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a2ebb49433142.12217517