Summary
Overall verdict: CLEAN (no evidence of fraud). This review examined the MICCAI 2023 paper introducing MedIM (DOI: 10.1007/978-3-031-43907-0_2) for potential academic misconduct. Three findings were central. First, the report format precluded pixel-level image analysis of Figures 1–3, so Western blot/PS-tampering checks could not be performed. Second, internal numerical consistency is strong: the last row of the ablation table (Table 3, full MedIM with L_align, KWM, and SDM) reports CheXpert, COVIDx, and SIIM scores of 89.25, 90.34, and 63.50, which match the Table 1 figures for the 10% labeled-data setting precisely—a signature unlikely under fabricated data. Third, reported AUC, accuracy, and recall values (e.g., 88.91, 77.22, 5.74) show natural last-digit variation consistent with genuine stochastic training. The timeline of datasets (MIMIC-CXR-JPG 2019, CheXpert 2019, COVIDx 2020, SIIM-ACR 2019) and baselines (MAE 2022, MGCA 2022, MRM 2023) predates the 2023 publication. Code is released. Confidence: moderate pending access to raw figures and code execution.
Verdict
CLEAN – No actionable evidence of academic fraud was identified. The paper passes internal-consistency, timeline, and plausibility checks. A formal image-forensics audit was not possible.
Key findings
- Pixel-level image analysis not performed: Figures 1–3 were not available in raw pixel form, so reuse, splicing, or PS manipulation in Western blots, gels, or microscopy images cannot be evaluated.
- Strong numerical consistency between main and ablation tables: The complete-model row in Table 3 (L_align + KWM + SDM) yields 89.25 (CheXpert), 90.34 (COVIDx), and 63.50 (SIIM), which exactly match the 10% labeled-data figures in Table 1—an alignment inconsistent with fabricated data.
- Plausible metric granularity: Reported AUC, accuracy, and recall values (e.g., 88.91, 77.22, 5.74) show natural last-digit variation rather than suspiciously round or monotonically progressing numbers; CheXpert performance scales plausibly from 1% to 100% labeled data.
- Chronologically valid references: Datasets (MIMIC-CXR-JPG 2019, CheXpert 2019, COVIDx 2020, SIIM-ACR 2019) and baselines (MAE 2022, MGCA 2022, MRM 2023) all predate the 2023 publication.
- Open code availability: The authors publicly release code, which is a positive credibility signal in computer-science research.
Evidence highlights
- DOI: 10.1007/978-3-031-43907-0_2
- Numerical match: Table 3 full-model row = Table 1 (10% labels) = 89.25 / 90.34 / 63.50 on CheXpert / COVIDx / SIIM.
- Sample metric values: 88.91, 77.22, 5.74 across Table 1 and Table 2.
- Datasets and baselines dated 2019–2023, all prior to publication year.
Notes
- The text-only nature of the review limits forensic checks of Fig. 1–3 and any supplementary images; a pixel-level re-examination is recommended if high-resolution originals become available.
- No statistical re-analysis (e.g., recomputation of variance, seed sensitivity) was performed; the natural last-digit variation is suggestive but not conclusive proof of authentic training runs.
- The review is AI-assisted and intended for academic discussion only; final determinations of misconduct require institutional investigation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a37ebc67b7a86.39832058