English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Geng Integrity Report: Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-training

Academic fraud report · Geng Detector

Summary

This report assesses the ACM Multimedia 2025 paper "Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-training" (Qiao et al., DOI: 10.1145/3746027.3755336) and rates it as highly suspicious. Four main concerns are documented. First, the prose in Section 5.3.3 cites Table 6 rows 1 and 3, and rows 2 and 3, but Row 1 is labeled Reconstruction only and Row 2 Alignment only, so the comparisons are mis-attributed by one row each. Second, the GLoRIA baseline (ResNet-50 backbone per its ICCV 2021 original) is placed in the ViT-based comparison group of Table 1 with no re-implementation note. Third, pretraining on MIMIC-CXR (377k images, 227k reports) with a multi-module model (ViT-B/16, BERT, cross-attention, GAT) for 200+15 epochs is claimed on only two RTX 4090 GPUs, which appears implausibly under-resourced. Fourth, Table 4 shows PAR+DKBA underperforming DKBA alone on ChestX-ray14 (80.8 vs 81.0 AUC), a negative result the authors do not discuss. Limitations: analysis is based on the published manuscript only; no raw logs or code were inspected, so final determination requires institutional investigation.

Verdict

Highly suspicious. Multiple internal inconsistencies and implausible methodological claims warrant formal investigation by the conference and the authors' institution.

Key findings

  • Table 6 row numbering contradicts the prose. The ablation discussion in Section 5.3.3 compares "joint" against "alignment-only" and "reconstruction-only," but cites rows that do not match the row labels.
  • GLoRIA is mis-categorized as ViT-based. Table 1 groups a ResNet-50 model under Vision Transformer baselines without justification.
  • Compute budget appears implausible. MIMIC-CXR pretraining with a ViT-B/16 + BERT + cross-attention + GAT model for 200 + 15 epochs on two RTX 4090 GPUs (24 GB each) is not credibly documented.
  • Selective reporting in Table 4. Adding PAR to DKBA reduces ChestX-ray14 AUC from 81.0 to 80.8; this negative result is not discussed.
  • Evidence highlights

  • Section 5.3.3 / Table 6: prose says "row 1 vs 3" for joint vs alignment-only, but Row 1 is Reconstruction only and Row 2 is Alignment only. The next sentence cites "row 2 vs 3" for joint vs reconstruction-only, again off by one.
  • Table 1: GLoRIA listed under ViT-based baselines; GLoRIA (ICCV 2021) is canonically ResNet-50 based.
  • Section 4.3: claims 2× NVIDIA RTX 4090 for pretraining on MIMIC-CXR (377k images, 227k reports); stage 1 = 200 epochs, stage 2 = 15 epochs.
  • Section 5.3.1 / Table 4: PAR + DKBA = 80.8 AUC vs DKBA only = 81.0 AUC on ChestX-ray14; not addressed in narrative.
  • Notes

  • All findings are derived solely from the published manuscript (DOI: 10.1145/3746027.3755336); no code, logs, or raw outputs were examined.
  • The row-numbering and GLoRIA-categorization errors could in principle be copy-editing or table-rendering mistakes, but their combination with the compute claim raises broader reproducibility concerns.
  • The compute estimate is qualitative; actual feasibility depends on batch size, precision, gradient accumulation, and implementation, none of which are fully disclosed.
  • Final determination of misconduct requires formal institutional investigation.

Tags

#academic-fraud#internal-inconsistency#image-manipulation#reproducibility#benchmark-misattribution#compute-feasibility#selective-reporting#chest-x-ray

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a366f0d5dfb88.48544006