English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Academic Integrity Review Report: Patient-Level Prediction of Multi-Classification Task at Prostate MRI Based on End-to-End Framework Learning From Diagnostic Logic of Radiologists

Academic fraud report · Geng Detector

Summary

This report assesses potential academic integrity concerns in a 2021 IEEE Transactions on Biomedical Engineering paper (DOI: 10.1109/TBME.2021.3082176) on end-to-end prostate MRI classification. The overall verdict is 'highly suspicious' (orange). Four issues were identified. The most severe finding concerns a statistical reporting anomaly: the Discussion claims significance of p<0.05 versus DCNN and radiomics baselines, but no p-values, test names, or statistical methodology appear anywhere in the results tables. A second key issue is an internal numerical contradiction in the dataset description (187 patients vs. N=178 in the same sentence, leaving 9 patients unaccounted for). A third issue involves suspiciously rounded and uniform accuracy values across independent external test sets, with TeS1 reporting exactly 0.800 and consistent standard deviation patterns. A fourth, lower-confidence finding notes unnatural oscillation in training curves and visually similar case-study images. Confidence is high for findings 1-3 based on the published text; finding 4 remains limited by lack of pixel-level evidence. Final determination requires institutional investigation.

Verdict

🟠 Highly suspicious. Multiple independent concerns are documented within the published text alone, with the most serious being a statistical-significance claim unsupported by any reported test or p-value. No allegation of fraud is made here; the findings are consistent with sloppy data handling, insufficient statistical reporting, or selective presentation.

Key findings

  • Statistical-significance claim without supporting evidence (Finding 3, 🔴): The Discussion states that the proposed method outperformed DCNN (p<0.05) and radiomics (p<0.05), yet Tables II, III, IV, and VIII contain no p-values and no statistical test is named anywhere in the methods.
  • Internal numerical contradiction in the dataset description (Finding 1, 🟠): Section IV.B refers to "187 patients of PUTH-p2 (N=178)" in a single clause, leaving 9 patients unaccounted for.
  • Suspiciously uniform and rounded accuracy figures (Finding 2, 🟠): Reported ACC-case values across independent external sets include 0.849±0.023 (TS), 0.824±0.051 (VS), 0.800±0.029 (TeS1), and 0.824±0.044 (TeS2); TeS1 lands exactly at 0.800 and standard-deviation patterns appear unusually regular.
  • Visual irregularity concerns (Finding 4, 🟡): Figure 6 training curves show unnatural sawtooth oscillation and the five "typical cases" in Figure 7 appear visually similar; flagged but not confirmed due to absence of pixel-level analysis.
  • Evidence highlights

  • Direct quote (Section IV.B, p. 3695): "Then, 187 patients of PUTH-p2 (N=178) without slice annotations are selected as an external test set (TeS1)." — The pair (187, N=178) is internally inconsistent.
  • Direct quote (Section VI, p. 3699): "In the experiment, our proposed method outperformed DCNN (p-value<0.05) and radiomics (p-value<0.05) significantly." — No corresponding p-values or statistical test procedures appear in Tables II–IV or VIII.
  • Quantitative pattern (Section V.B, Table III): ACC-case = 0.849±0.023 (TS); 0.824±0.051 (VS); 0.800±0.029 (TeS1); 0.824±0.044 (TeS2). TeS1 mean equals exactly 0.800, an unusually clean value for cross-center medical AI evaluation.
  • DOI: 10.1109/TBME.2021.3082176 — retained verbatim for traceability.
  • Notes

  • Findings 1–3 are derived solely from the published PDF text and tables and are verifiable by any reader; confidence is high.
  • Finding 4 relies on visual-summary heuristics applied to figures without high-resolution pixel inspection and is therefore marked as evidence-insufficient; it should not be treated as a confirmed irregularity.
  • The dataset discrepancy (187 vs. N=178) could reflect a typographical error, an inclusion/exclusion criterion not described, or a mislabeled subset; clarification from the authors is required.
  • The p-value claim may indicate that tests were performed but omitted from the manuscript — a remediable reporting defect — or that the claim is unsupported, which would be more serious.
  • No determination of misconduct is made here. Institutional review and author response (raw data, statistical scripts) are necessary before any formal conclusion.

Tags

#academic-integrity#medical-imaging#prostate-MRI#statistical-reporting#data-consistency#deep-learning#IEEE-TBME#review-report

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a674eafd42a89.60964541