English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Geng integrity review: "Anomaly detection on MVTec AD using VQ-VAE-2" (Procedia CIRP 130, 2024)

Academic fraud report · Geng Detector

Summary

This review assesses a 2024 conference paper by Edward K. Y. Yapp and Ngoc C. N. Doan comparing a VQ-VAE-2 model against single-level VQ-VAE variants on the MVTec AD benchmark. The central finding is highly suspicious tabular data in Table 3 (Page 1813): for the Capsule, Pill, Screw, and Carpet categories, the three compared models report image-level F1 scores that are identical to three decimal places (e.g., Capsule: 0.905/0.905/0.905; Pill: 0.916/0.916/0.916; Screw: 0.853/0.853/0.853; Carpet: 0.864/0.864/0.864). Producing bit-identical F1 scores from architecturally distinct models on the same dataset is statistically implausible, suggesting either copy-paste errors during data transfer into LaTeX or fabricated numbers. The textual claim that "VQ-VAE-2 is superior to VQ-VAE across all the metrics" is directly contradicted by these rows. The figure analysis could not be fully verified due to garbled PDF text extraction in Fig. 2, and the reported timeline (data gathered December 2023, conference 2024) is internally consistent. Overall verdict: highly suspicious (🟠), pending GitHub code reproduction and raw logs.

Verdict

🟠 Highly suspicious. The main red flag is Table 3, where four MVTec AD categories show bit-identical image-level F1 scores across three different models, contradicting the paper's narrative claim of VQ-VAE-2 superiority. The anomaly is consistent with either a copy-paste error during table generation or fabricated data.

Key findings

  • Identical F1 scores across distinct models (Table 3, p. 1813): Capsule = 0.905/0.905/0.905; Pill = 0.916/0.916/0.916; Screw = 0.853/0.853/0.853; Carpet = 0.864/0.864/0.864. Identical to three decimal places between VQ-VAE-2 and the VQ-VAE (top) and VQ-VAE (bottom) baselines.
  • Narrative–data contradiction (Section 4, pp. 1812–1813): The paper states "VQ-VAE-2 is superior to VQ-VAE across all the metrics," but the four identical rows provide no evidence of superiority, undermining the abstract's "statistical improvement" claim.
  • Figure inspection inconclusive (Fig. 2): PDF text extraction produced garbled output (e.g., !$ $ $ $), preventing pixel-level verification of the heatmaps. Other figures are vector architecture diagrams or statistical plots and show no obvious splicing from text-level review.
  • Timeline consistent (Footnote 2): CFA results dated 5 December 2023; conference and publication in 2024; methodology (PyTorch, Tesla V100) is plausible. No temporal contradictions found.
  • Code availability check pending: A GitHub repository (https://github.com/edwardyapp/vq-vae-2-pytorch-ad) is cited but has not been executed to reproduce Table 3.
  • Evidence highlights

  • Table 3, p. 1813: Four categories exhibit identical three-decimal F1 values across three model columns — extremely unlikely for distinct VQ-VAE-2 / VQ-VAE (top) / VQ-VAE (bottom) variants on the same MVTec AD split.
  • Section 4 prose vs. Table 3: Author asserts VQ-VAE-2 superiority on "all metrics," yet the identical rows show no measurable difference.
  • Fig. 2: Unreadable extracted text prevents visual forensic verification; flagged as unresolved.
  • Notes

  • DOI: 10.1016/j.procir.2024.10.320
  • Recommended follow-ups: (1) clone and run the cited GitHub code to attempt reproduction of Table 3; (2) request raw CSV/log files from the authors for the suspect categories; (3) raise a PubPeer comment requesting an erratum or formal explanation for the Capsule, Pill, Screw, and Carpet rows; (4) notify *Procedia CIRP* editors of the data anomaly.
  • Confidence is moderate-to-high on the tabular anomaly, low on figure integrity due to extraction failure. Final determination requires author-supplied raw outputs and independent code reproduction.
  • This report is AI-assisted and intended for academic discussion only; definitive misconduct findings must come from a formal institutional or editorial investigation.

Tags

#academic-fraud#data-fabrication#copy-paste-error#table-inconsistency#image-manipulation-pending#procedia-cirp#mvt-ec-ad#vq-vae

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a1e93fc717ce7.11330406