Summary
This report assesses a 2025 ACM Multimedia (MM '25) submission by Lihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao, Weisheng Li, and Xinbo Gao on chest X-ray vision-language pre-training. Overall verdict: highly suspicious. Three independent textual and tabular inconsistencies were identified. (1) In Table 1 (Linear classification), the authors' ViT-based method reports identical CXP AUC of 89.5 for both 1% and 10% training data, an implausible coincidence given a tenfold increase in training size. (2) The text claims a 2.0% AUC improvement over MRM on CXR14 at 1% data, but Table 1 shows 81.4 minus 78.8 equals 2.6, a clear numerical mismatch. (3) Table 1 contains a formatting artifact in the Med-UniC row where digits appear concatenated (e.g., '90.891.993.1'), suggesting insufficient quality control. Image-based forensic checks (PS traces, image reuse) were not applicable. Confidence is moderate-to-high for findings 1-3, limited only by lack of raw training logs.
Verdict
Highly suspicious (🟠). Three textual-tabular inconsistencies in a single ACM MM '25 submission suggest careless fabrication or copy-paste errors that exceed normal typographic noise.
Key findings
- Identical AUC values across 10× data increase: In Table 1, the proposed method's ViT-based CXP AUC is reported as 89.5 for both 1% and 10% training data, an outcome that is statistically implausible and indicative of data copying.
- Numerical mismatch between prose and Table 1: The paper states a 2.0% AUC improvement over MRM on ChestX-ray14 (1% data), yet Table 1 shows 81.4 vs 78.8, a difference of 2.6 percentage points.
- Formatting/data-parsing artifact: The Med-UniC ViT-based row in Table 1 shows concatenated digits ("90.891.993.193.780.389.594.5"), indicating broken delimiter handling and poor submission hygiene.
- Image-based forensics not applicable: The paper contains only architecture diagrams, t-SNE plots, and Grad-CAM heatmaps; no Western blots, microscopy, or gels are present, so pixel-level reuse checks were waived.
Evidence highlights
- Table 1, ViT-based section, CXP column:
1% and 10% both listed as 89.5.
- Section 5.1, first paragraph: "On the ChestX-ray14 dataset, using only 1% of the data for fine-tuning, our framework achieves a 2.0% AUC score improvement over MRM."
- Table 1 CXR14 1% column: MRM = 78.8, Ours = 81.4 (Δ = 2.6, not 2.0).
- Table 1, Med-UniC ViT-based row exhibits run-together numeric tokens.
- DOI: 10.1145/3746027.3755336.
Notes
- All findings are derived from the text and tables of the submission PDF; raw training logs, random seeds, and code were not available.
- The identical-AUC anomaly is necessary but not sufficient to prove fabrication; benign explanations (e.g., post-hoc rounding from a saturated model) cannot be fully ruled out without logs.
- The 2.0% vs 2.6% discrepancy is a deterministic arithmetic conflict that is hard to attribute to rounding.
- Image-manipulation, western-blot reuse, and gel duplication checks were not applicable for this computer-science submission.
- Recommended actions: request full training logs and random seeds from the authors, query the 2.0% vs 2.6% mismatch, post a PubPeer comment, and have the Area Chair verify the released code repository.
- This report is AI-assisted and intended for academic discussion only; final determinations of misconduct require formal institutional investigation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a37730c582c70.90751197