Verdict
Highly suspicious. The paper contains multiple apparent reporting and internal-consistency anomalies. The duplicate CheXpert AUC of 89.5 for 1% and 10% training data, together with the contradiction between Section 5.1 and Table 1, is the strongest evidence supporting closer scrutiny. These issues do not independently prove fabrication. The concerns about two NVIDIA RTX 4090 GPUs completing the stated MIMIC-CXR pretraining in an implausible time remain under-supported without information on batch size, software stack, optimization settings, and measured runtime. A formal conclusion requires raw outputs, logs, code, and an official editorial investigation.
Key findings
- Identical CheXpert results: In Table 1, under the ViT-based setting, the proposed model reportedly obtains an AUC of 89.5 with both 1% and 10% training data. This is a direct numeric inconsistency that requires explanation.
- Text–table contradiction: Section 5.1 states that the proposed model underperforms Med-UniC on CheXpert. Nevertheless, Table 1 reports 89.5 for the proposed model and 89.4 for Med-UniC in the 1% ViT-based setting.
- Selective ablation interpretation: Adding DKBA alone is associated with a decline from 64.9 to 64.5 on SIIM and from 20.2 to 19.9 in RSNA mAP. The report finds that Section 5.3.1 focuses only on positive outcomes and does not discuss these negative results.
- Unverified feasibility concern: Pretraining on more than 370,000 MIMIC-CXR images for 200 epochs using two NVIDIA RTX 4090 GPUs may be computationally demanding. The available information is insufficient to conclude that the reported training was impossible or exaggerated.
- DOI: 10.1145/3746027.3755336
- Table 1, CXP, ViT-based:
- Proposed model at 1%: 89.5
- Proposed model at 10%: 89.5
- Med-UniC at 1%: 89.4
- Table 4 and Section 5.3.1:
- SIIM: 64.9 to 64.5
- RSNA mAP: 20.2 to 19.9
- The 4090-related concern concerns implementation feasibility, not an established computational impossibility.
Evidence highlights
Notes
The duplicated AUC may potentially arise from transcription, table-entry, or figure-generation errors, but the supplied report does not provide original laboratory records that would distinguish error from misconduct. The mismatch between the narrative and Table 1 may similarly result from revisions or incorrect table assembly. Nevertheless, the combined anomalies warrant verification of the underlying data and manuscript provenance. Request the complete MIMIC-CXR training logs, TensorBoard or Weights & Biases records, data-split definitions, random seeds, raw evaluation outputs, and reproduction scripts. Editorial review should also determine whether the negative DKBA ablation results were considered and how they affected the paper’s interpretation. Any final finding of fabrication or academic misconduct should be based on an institutional or conference investigation rather than this report alone.