English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Investigation Report: "ChatGPT Combining Machine Learning for the Prediction of Nanozyme Catalytic Types and Activities" (DOI: 10.1021/acs.jcim.4c00600)

Academic fraud report · Geng Detector

Summary

This report assesses a 2024 publication in the Journal of Chemical Information and Modeling that combines ChatGPT-based data extraction with machine learning to predict nanozyme catalytic activity. The overall verdict is highly suspicious (orange). Key concerns include severe overfitting in the Gradient Boosting Regressor (GBR) model: training R² for Km reached 0.9985 while test R² dropped to 0.6476 (Km) and 0.6834 (Vmax), and MAE rose from 0.0068 to 0.5324 (~80× increase). Authors selectively highlighted the stronger Kcat test result (R²=0.95) while downplaying weaker targets. An external validation used an extraordinarily lenient threshold of |Δlg Km| ≤ 1, meaning predicted values within a 10-fold range of true values were counted as correct, yielding an inflated 81% accuracy. ChatGPT extraction agreement with human curators was only 65.11% on 102 test papers, requiring manual rescue and k-NN imputation. The Nanozyme Copilot evaluation used only n=10 synthesis pathways to claim >90% accuracy. Confidence is moderate; findings are based on reported metrics in the published paper and were not independently re-run.

Verdict

🟠 Highly suspicious. The paper reports multiple indicators of overstated claims, severe overfitting, and questionable validation methodology. Concerns are serious enough to warrant editorial review.

Key findings

  • Severe train/test performance gap (overfitting). Reported GBR metrics show Km training R² = 0.9985 vs. test R² = 0.6476; Vmax test R² = 0.6834. Test MAE rose from training 0.0068 to 0.5324, an ~80× increase, inconsistent with the paper's claim of "superior prediction performance."
  • Selective reporting. The abstract emphasizes the Kcat test result (R² = 0.95) while obscuring weak Km and Vmax test performance.
  • Implausibly lenient external-validation threshold. Accuracy on external literature was computed using |Δlg Km| ≤ 1 or |Δlg Vmax| ≤ 1, which on log-transformed data corresponds to a 10-fold tolerance between predicted and true values. The reported 81% accuracy under this criterion is not meaningful as a performance benchmark.
  • Noisy data extraction pipeline. ChatGPT agreement with human extraction was 65.11% on 102 papers. A second-round human pass plus k-NN imputation were required, raising concerns about the integrity of training data.
  • Statistically inadequate Copilot evaluation. The Nanozyme Copilot was assessed on n = 10 randomly selected synthesis pathways to support a >90% accuracy claim; the sample size is far too small for any defensible conclusion.
  • Padding suspicion (lower confidence). Equations 1–8 derive standard AdaBoost weight updates and define R² and MAE; this is flagged as filler content but evidence of intent is inconclusive.
  • Evidence highlights

  • DOI: 10.1021/acs.jcim.4c00600
  • Reported metrics: Km train R² = 0.9985; Km test R² = 0.6476; Vmax test R² = 0.6834; MAE train = 0.0068 → test = 0.5324 (≈80× increase).
  • ChatGPT–human extraction agreement = 65.11% (102-paper test).
  • External-validation accuracy = 81% using |Δlg| ≤ 1 (equivalent to 10× tolerance on original scale).
  • Nanozyme Copilot evaluation sample size: n = 10.
  • Notes

  • All numeric values above are reproduced from the published paper as cited; no independent reanalysis was performed.
  • Severity ratings and the overall verdict reflect the patterns of selective reporting, overfitting, and methodologically weak validation rather than a finding of confirmed misconduct.
  • Recommended follow-up: request raw training/test splits and extraction logs; raise concerns on PubPeer regarding the test R² collapse and the |Δlg| ≤ 1 tolerance; request that the journal reassess the abstract's "superior prediction performance" wording.
  • This is an AI-assisted integrity review and is not an official determination of misconduct.

Tags

#academic-fraud#overfitting#selective-reporting#machine-learning#data-quality#validation-methodology#chatgpt#nanozyme

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a421b8ea64962.39045305