Verdict
Questionable, with no affirmative evidence of academic misconduct identified. The assessment covers reported Tables 1–6, implementation details, references, and the captions for Figures 1–6. It does not establish fabrication, data falsification, statistical manipulation, or image duplication. However, the original figures were unavailable, so image reuse and splicing could not be excluded through pixel-level analysis.
Key findings
- The evaluation could not perform background-noise, Western blot, edge-cropping, or other visual analyses because only textual material and figure captions were available.
- No anomalous, contradictory, or evidently duplicated figure captions were identified for Figures 1, 2, 3, and 5.
- The reported results contain non-monotonic and adverse ablation outcomes rather than uniformly improved scores.
- Table 4 reportedly gives a standalone PAR result of 78.9 on ChestX-ray14, below the no-module baseline of 79.1.
- Standalone DKBA reportedly yields 19.9 mAP on the RSNA detection task, below the corresponding baseline of 20.2.
- Table 1 reportedly shows the authors’ method at 89.7 on CheXpert, below Med-UniC at 90.8; the authors attribute this to Med-UniC’s use of additional pretraining data.
- The report found no arithmetic progression, repeated trailing-digit pattern, or obvious copy-and-paste pattern in the reported metrics.
- The stated use of 2 NVIDIA RTX 4090 GPUs for ViT-B/16 pretraining was considered reasonable rather than implausibly excessive.
- The reported citation of 2024 work in a 2025 publication was judged temporally coherent.
- The GitHub link
https://github.com/Felix1118/PADKBis cited in the abstract, but its authenticity and functionality were not independently verified.
Evidence highlights
The principal evidence supporting the absence of obvious statistical red flags is the reported presence of “imperfect” experimental results: 78.9 for PAR alone versus 79.1 for the no-module baseline on ChestX-ray14, and 19.9 for DKBA alone versus 20.2 for the baseline on the RSNA detection task. The CheXpert comparison is also characterized by a lower reported score of 89.7, compared with 90.8 for Med-UniC. These patterns are consistent with the report’s view that the tables reflect real module-specific trade-offs rather than uniformly inflated gains.
The review additionally found the hardware description—2 NVIDIA RTX 4090 GPUs—plausible for the stated model and dataset scale. It also found no apparent temporal conflict involving the 2024 Med-UniC reference in a 2025 ACM MM paper. Nevertheless, these textual consistencies are not proof that the underlying experiments or visual assets are authentic.
Notes
The review explicitly states that Figures 1–6 were not available for original-image or pixel-level comparison. Therefore, conclusions about image reuse, image splicing, background inconsistencies, or other visual manipulation remain unresolved. The report recommends examining the released code and reproducing the experiments where feasible. No specific author contact, PubPeer submission, or editorial complaint was identified as a necessary action on the evidence presented. Any formal finding of academic misconduct would require investigation by qualified institutions and access to the original figures, data, code, and complete experimental records. The cited paper is “Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-training,” DOI: 10.1145/3746027.3755336.