Summary
This report catalogs multiple serious anomalies in Junjie Jiang et al.'s paper (DOI: 10.3390/electronics14244785), published in Electronics (MDPI) and accepted on 26 November 2025. The verdict is strongly adverse based on five independent discrepancies. The most severe findings are: (1) a temporal impossibility—an acceptance date of 26 November 2025 with multiple references dated to 2026 (e.g., Expert Syst. Appl. 2026, 297, 129489); (2) an inconsistent parameter count of 'approximately 436 million' for a dual-backbone architecture that combines CLIP ViT-L/14@336px (~427M) and DINOv2 ViT-L (~304M), which should exceed 730M parameters; and (3) a self-contradictory architecture description that pairs DINOv2 ViT-L with '12 blocks,' a configuration belonging to ViT-B, not ViT-L. Additional concerns include cherry-picked comparative metrics, suspiciously uniform 0.1–0.2% gains over the Crane baseline, and an implausibly brief revision window of 2 days. Confidence is high on the future-dated references and parameter-count contradiction, which are verifiable from the published text alone. Limitations: we could not independently verify code or raw data.
Verdict
Strong adverse verdict (实锤). The paper contains at least three independently verifiable factual errors of high severity—futures-dated references, an arithmetically impossible parameter count, and a ViT-L/ViT-B architectural contradiction—along with two moderate-severity concerns regarding data presentation and review timeline.
Key findings
- Future-dated references (🔴 severe): Accepted 26 November 2025 yet cites at least four 2026 publications, including Expert Syst. Appl. 2026, 297, 129489 and Adv. Eng. Inform. 2026, 69 Pt B, 103931.
- Impossible parameter count (🔴 severe): Claims "approximately 436 million parameters" for a CLIP ViT-L/14@336px + DINOv2 ViT-L dual backbone. CLIP ViT-L total is ~427M; DINOv2 ViT-L is ~304M; combined should be ≥730M, not 436M.
- Architectural self-contradiction (🔴 severe): Section 3.3 states DINOv2 is used "with 12 blocks" while also naming DINOv2 ViT-L, which has 24 blocks. ViT-B has 12 blocks, suggesting possible misrepresentation of model size.
- Selective reporting (🟠 moderate): Pixel-level AUPRO on MVTec in Table 5 shows Ours (87.0) below Crane (88.1), contrary to the textual claim of consistent superiority over Crane.
- Implausibly uniform gains (🟠 moderate): On KSDD, DAGM and others, margins over Crane cluster at 0.1%–0.2%, an unnatural distribution suggestive of data tuning.
- Abnormal review timeline (🟠 moderate): Received 28 October 2025; revised 24 November 2025; accepted 26 November 2025—2 days from revision to acceptance, atypical for a deep-learning study covering 7 datasets with ablations.
Evidence highlights
- DOI: 10.3390/electronics14244785
- Refs. [1], [3], [6] in *Expert Syst. Appl.* 2026, 296/297; Ref. [7] in *Adv. Eng. Inform.* 2026, 69 Pt B, 103931—impossible given the 2025 acceptance date.
- Section 4.1.3 reports "approximately 436 million parameters" while specifying CLIP ViT-L/14@336px + DINOv2 ViT-L, whose known sizes total ≥730M.
- Section 3.3: "The DINOv2 model (with 12 blocks)… The ViT-L model (with 24 blocks)…"—standard ViT-L has 24 blocks, not 12.
- Table 5 (MVTec, Pixel-AUPRO): Ours 87.0 < Crane 88.1, contradicting Section 4.2 narrative.
- Timeline: Received 28 Oct 2025; Revised 24 Nov 2025; Accepted 26 Nov 2025 (2-day revision-to-acceptance window).
Notes
- Document type: AI-assisted investigative review; final determination requires institutional investigation.
- The future-dated references and the 436M vs ≥730M arithmetic discrepancy are directly verifiable from the PDF text and require no code access.
- The ViT-L vs 12-blocks inconsistency is a textual fact but should be cross-checked against any released source code or weights.
- Pixel-AUPRO underperformance on MVTec does not, by itself, prove misconduct; it indicates selective narrative framing.
- Report preserves the original finding IDs and severities translated from the Chinese source.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a1e4b3d149503.89659887