English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Geng Integrity Report: Contrastive Masked Image-Text Modeling for Medical Visual Representation Learning (MICCAI 2023, DOI: 10.1007/978-3-031-43904-9_48)

Academic fraud report · Geng Detector

Summary

Verdict: 🟡 Questionable. The report flags the CMITM paper (Chen et al., MICCAI 2023, LNCS vol. 14224) for lacking statistical rigor and for narrative claims that are partially contradicted by its own tables. Key issues: (1) The authors describe 'consistent improvements' over baselines, but several gaps in Table 1 (p. 499) are only 0.1–0.2 (e.g., CheXpert 100% labels: CMITM 89.2 vs MRM 88.7), well within typical deep-learning run-to-run variance, yet no mean±SD across multiple seeds or significance tests are reported. (2) The claim of general superiority conflicts with the data: on RSNA with 10% labels CMITM scores 92.6 vs MRM 92.7, and on COVIDx with 10% labels both score 90.2. (3) No image-based forensics could be performed because figures are algorithmically generated and pixel data was not supplied. Confidence: moderate; conclusions rest on the reported numbers and absent statistics, not on independent re-implementation.

Verdict

🟡 Questionable. The paper contains internal inconsistencies between its narrative claims and tabular results, and lacks the statistical apparatus (multiple seeds, error bars, significance tests) needed to substantiate marginal differences in deep-learning benchmarks. No image-manipulation evidence was assessable.

Key findings

  • Marginal numeric gains presented as 'consistent improvements': gaps of 0.1–0.2 percentage points are reported without mean±SD or p-values.
  • Internal contradiction: text claims CMITM 'generally outperforms' masked-autoencoding or contrastive-only baselines, yet Table 1 shows a loss on RSNA@10% (92.6 vs 92.7) and a tie on COVIDx@10% (90.2 vs 90.2).
  • No reproducibility safeguards disclosed in the paper itself: no multiple-seed averaging, no confidence intervals, no significance testing.
  • Image-based integrity checks not feasible: Figures 1–3 are schematic/training plots; pixel-level analysis requires the original images, which were not available.
  • Authors state code is released at https://github.com/cchen-cc/CMITM, which mitigates—but does not eliminate—the risk of selective-result reporting.
  • Evidence highlights

  • Table 1 (p. 499), CheXpert 100% labels: CMITM 89.2 vs MRM 88.7 (Δ ≈ 0.5, within likely seed-driven variance).
  • Table 1 (p. 499), RSNA 10% labels: CMITM 92.6 vs MRM 92.7 (CMITM lower).
  • Table 1 (p. 499), COVIDx 10% labels: CMITM 90.2 vs MRM 90.2 (tie).
  • Absence of error bars, standard deviations, or p-values across all reported numbers.
  • DOI: 10.1007/978-3-031-43904-9_48; venue: MICCAI 2023, Lecture Notes in Computer Science vol. 14224.
  • Notes

  • Limitations of this report: only the textual content of the PDF was reviewed; raw figure pixels and the published code were not independently executed. The 'consistent improvements' language and the absence of statistical tests are concerning but, on their own, do not constitute proof of misconduct. Recommended follow-up: (1) request multi-seed results and significance tests from the authors on PubPeer; (2) reproduce the GitHub release and compare against reported numbers; (3) request training logs (TensorBoard/WandB) to rule out cherry-picking. All findings are based on the paper as cited; any updated version (e.g., with errata or supplementary statistics) may alter this assessment.

Tags

#academic-integrity#deep-learning#medical-imaging#missing-statistics#internal-inconsistency#questionable-claims#miccai-2023#reproducibility

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a37df21d7d854.77956520