Verdict
🟠 Highly Suspicious. Two independent, text-verifiable data anomalies have been confirmed in the reported results. No image-level analysis could be performed from the available source.
Key findings
- Finding 1 (Confirmed, Severity 🔴): Identical AUC across 1% and 10% training data. In Table 1 (p.5), CheXpert (CXP) ViT-based results for the proposed method are reported as 89.5 at both 1% and 10% training data. The baseline MRM likewise shows 88.5 at both ratios. Two distinct models producing identical, one-decimal AUC values at two different data scales is statistically implausible and indicative of manual fabrication or copy-paste error.
- Finding 2 (Confirmed, Severity 🔴): Ablation contradicts stated conclusion. Table 4 (p.6) reports ChestX-ray14 AUC of 81.0 for DKBA alone, but 80.8 for PAR + DKBA (full proposed method). Section 5.3.1 nevertheless asserts "complementary benefits" of PAR and DKBA. A drop of 0.2 AUC from adding a module is logically incompatible with a complementarity claim.
- Finding 3 (Unverified, Severity 🟠): Compute scale plausibility concern. Section 4.3 reports pre-training for 200 epochs (reconstruction) plus 15 epochs (alignment) on MIMIC-CXR (377,110 images) with ViT-B/16, BERT, and GAT components, on only two NVIDIA RTX 4090 GPUs. Whether this is feasible within the stated timeline was not independently confirmed.
- Finding 4 (Unverified, Severity 🟡): Image reuse / splicing. Pixel-level inspection of Figures 1–6 was not possible; t-SNE and qualitative figure comparisons are recommended if the PDF is obtained.
- Table 1 (p.5), CXP ViT-based, Proposed method: AUC 1% = 89.5; AUC 10% = 89.5.
- Table 1 (p.5), CXP ViT-based, MRM baseline: AUC 1% = 88.5; AUC 10% = 88.5.
- Table 4 (p.6), ChestX-ray14 ablation: DKBA = 81.0; PAR + DKBA = 80.8.
- Section 4.3, Implementation Details: two NVIDIA RTX 4090 GPUs; 200 reconstruction pre-training epochs + 15 alignment epochs on MIMIC-CXR (377,110 images).
- Section 5.3.1, narrative claim: PAR and DKBA exhibit "complementary benefits," contradicting the observed ablation drop.
- DOI of subject paper: 10.1145/3746027.3755336.
- Findings 1 and 2 are based on the numbers as printed and are reproducible from the paper text; no external datasets or image processing were used.
- Finding 3 is a plausibility concern only; RTX 4090 (24 GB VRAM each) is not, on its face, impossible for reduced batch sizes or gradient accumulation, so this is not treated as evidence of fabrication.
- Finding 4 cannot be evaluated without access to the original PDF and figures.
- All numeric values are quoted exactly as they appear in the source report.
- Recommended next actions: request raw training logs and per-run seed results from the authors; raise the Table 1 / Table 4 inconsistencies on PubPeer; forward concerns to the ACM MM '25 program committee and the authors' institutions if no satisfactory explanation is provided.