Summary
This report raises four substantive concerns against the above Nature Communications paper, with an overall verdict of highly suspicious. Finding 1 documents an institutional name error: Harvard Medical School is rendered as "Harvard Medical University," an entity that does not exist. Finding 2 identifies an arithmetic impossibility in the RSNA Pneumonia Detection dataset description, where 25,184 + 1,500 + 3,000 = 29,684 exceeds the official ~26,684 images by 3,000. Finding 3 highlights implausibly large performance jumps on COVID Rural segmentation (MaCo at 75.1% Dice vs. the previous best of 44.0% MedKLIP at 100% labels, a >70% relative gain) without adequate ablations. Finding 4 flags implausible compute claims: 3.5 hours on four A100 GPUs to pretrain on 377,000 image-report pairs, which is inconsistent with known training times for ViT/BERT-scale multimodal models. No image manipulation was assessed due to lack of figures. The findings are based on text, arithmetic and logical analysis; final determination of misconduct requires institutional investigation.
Verdict
Highly suspicious. Multiple basic factual, arithmetic, and computational-plausibility issues were identified, none of which appear individually minor; together they suggest either fabrication, extreme carelessness, or outsourced writing without subject-matter oversight.
Key findings
- Non-existent institutional affiliation (severity: critical). The author Hong-Yu Zhou is listed under "Department of Biomedical Informatics, Harvard Medical University, Boston, MA, USA." No such institution exists; Harvard's medical school is Harvard Medical School (HMS).
- Dataset arithmetic inconsistent with public record (severity: critical). RSNA Pneumonia Detection is reported as split into 25,184 training + 1,500 validation + 3,000 test images, totaling 29,684, whereas the official dataset contains approximately 26,684 images — a discrepancy of 3,000 images.
- Implausible performance gain (severity: high). On COVID Rural segmentation, MaCo reports 75.1% Dice versus the prior best of 44.0% (MedKLIP) at 100% labels, a >30 absolute-point / >70% relative improvement over eight SOTA baselines, discussed in a single sentence without ablations or failure-case analysis.
- Implausible training cost claim (severity: high). Pre-training on >377,000 image-report pairs of MIMIC-CXR V2 with a ViT-B/16 image encoder, BERT text encoder, masked autoencoding and contrastive objectives is stated to take ~3.5 hours on four NVIDIA A100 GPUs, which is inconsistent with reported training durations for comparable multimodal foundation models.
Evidence highlights
- Affiliation text: "Department of Biomedical Informatics, Harvard Medical University, Boston, MA, USA." (Page 1 footnote). Harvard Medical University is not a recognized institution; Harvard Medical School (HMS) is the correct name.
- Dataset numbers (Page 9, Methods — Datasets, RSNA Pneumonia Detection paragraph): "training set of 25,184 images, a validation set of 1500 images, and a test set of 3,000 images." Sum = 29,684 vs. official ~26,684.
- Table 3 (Page 4, fine-tuning segmentation): COVID Rural 100% labels — MaCo Dice 75.1 vs. MedKLIP 44.0; previous SOTA range cited across eight methods.
- Implementation details (Page 8): "The pre-training of MaCo was completed in approximately 3.5 hours using four NVIDIA A100 GPUs" on MIMIC-CXR V2 with >377,000 image-report pairs.
- DOI: 10.1038/s41467-024-51749-0
Notes
- Scope: this review is limited to text, arithmetic and logical plausibility; no image-based duplicate or splicing analysis was possible because figures were not provided for pixel inspection.
- The institutional-name error and the RSNA count error are independently verifiable and, if confirmed against the published version, would constitute basic factual mistakes of a kind not expected from authors with genuine primary access to the data and institutions claimed.
- The two high-severity findings (segmentation jump and training-time claim) are indicators rather than proof; they require re-running experiments with disclosed code, seeds and logs.
- Authors and editors have not yet been contacted as part of this report; recommended follow-ups include requesting raw splits and training logs, posting on PubPeer, notifying the Nature Communications editorial office, and verifying Zhou's affiliation with Harvard Medical School's Office of Academic Affairs.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a37bcde7058b5.49424214