Verdict
🟠 Highly suspicious. The paper reports multiple indicators of overstated claims, severe overfitting, and questionable validation methodology. Concerns are serious enough to warrant editorial review.
Key findings
- Severe train/test performance gap (overfitting). Reported GBR metrics show Km training R² = 0.9985 vs. test R² = 0.6476; Vmax test R² = 0.6834. Test MAE rose from training 0.0068 to 0.5324, an ~80× increase, inconsistent with the paper's claim of "superior prediction performance."
- Selective reporting. The abstract emphasizes the Kcat test result (R² = 0.95) while obscuring weak Km and Vmax test performance.
- Implausibly lenient external-validation threshold. Accuracy on external literature was computed using |Δlg Km| ≤ 1 or |Δlg Vmax| ≤ 1, which on log-transformed data corresponds to a 10-fold tolerance between predicted and true values. The reported 81% accuracy under this criterion is not meaningful as a performance benchmark.
- Noisy data extraction pipeline. ChatGPT agreement with human extraction was 65.11% on 102 papers. A second-round human pass plus k-NN imputation were required, raising concerns about the integrity of training data.
- Statistically inadequate Copilot evaluation. The Nanozyme Copilot was assessed on n = 10 randomly selected synthesis pathways to support a >90% accuracy claim; the sample size is far too small for any defensible conclusion.
- Padding suspicion (lower confidence). Equations 1–8 derive standard AdaBoost weight updates and define R² and MAE; this is flagged as filler content but evidence of intent is inconclusive.
- DOI: 10.1021/acs.jcim.4c00600
- Reported metrics: Km train R² = 0.9985; Km test R² = 0.6476; Vmax test R² = 0.6834; MAE train = 0.0068 → test = 0.5324 (≈80× increase).
- ChatGPT–human extraction agreement = 65.11% (102-paper test).
- External-validation accuracy = 81% using |Δlg| ≤ 1 (equivalent to 10× tolerance on original scale).
- Nanozyme Copilot evaluation sample size: n = 10.
- All numeric values above are reproduced from the published paper as cited; no independent reanalysis was performed.
- Severity ratings and the overall verdict reflect the patterns of selective reporting, overfitting, and methodologically weak validation rather than a finding of confirmed misconduct.
- Recommended follow-up: request raw training/test splits and extraction logs; raise concerns on PubPeer regarding the test R² collapse and the |Δlg| ≤ 1 tolerance; request that the journal reassess the abstract's "superior prediction performance" wording.
- This is an AI-assisted integrity review and is not an official determination of misconduct.