Summary
This report assesses a 2026 Analytical Chemistry paper by Weixiang Huang et al. proposing a cascaded neural network (CSAM-ResUNet) for Raman spectral analysis of mixed microplastics. The overall verdict is highly suspicious. The principal concern is in Table 6, where validation and test metrics for spectral unmixing are nearly identical to five decimal places (e.g., Component 1: val/test R² = 0.99959, MSE = 9.20×10⁻⁷; Component 2: R² = 0.99955), which is statistically implausible for independent splits and strongly suggests copy-paste authoring or a non-independent data partition. A secondary concern is that in the lowest-energy Dataset 9 (50 mW, 1000 ms; 50 mJ total), the raw classification accuracy is only 8.97%—near random—yet the model purportedly achieves 82.51%, which is physically implausible and indicates overfitting on synthetic data or hallucinated mappings. PDF text extraction issues (e.g., '0.999628.60×10⁻⁷') limited confirmation of finer digit distributions. Confidence is high for Finding 1 and moderate for Finding 2; full confirmation requires original datasets, training logs, and independent test scripts.
Verdict
🟠
Highly suspicious. Two independent red flags—implausible metric duplication and physically implausible low-SNR recovery—warrant formal investigation by the journal and the authors' institution.
Key findings
- Independent-split metrics collapse in Table 6 (Page 8020). Validation and test unmixing metrics are reported to identical precision (5 decimal places for R², identical MSE values) for both components. This is statistically improbable for independent evaluation sets and points to either copy-paste table construction or non-independent partitioning.
- Physically implausible low-energy classification in Table 4 (Page 8017). Under Dataset 9 (50 mW, 1000 ms; 50 mJ total), raw accuracy is 8.97% (random-guess level), yet the proposed model claims 82.51%. A neural network cannot recover chemical identity from spectra whose features are buried in noise.
- Text-extraction artifacts compromise finer forensic checks. Numerals and scientific notation are concatenated in the PDF (e.g.,
0.999628.60×10⁻⁷), making column-level verification of Tables 1 and 2 unreliable. Evidence highlights
- Table 6, Component 1: validation R² = 0.99959, test R² = 0.99959; validation MSE = 9.20×10⁻⁷, test MSE = 9.20×10⁻⁷.
- Table 6, Component 2: validation R² = 0.99955, test R² = 0.99955 (to 5 decimal places).
- Table 4, Dataset 9 (50 mW, 1000 ms): raw accuracy 8.97% → CSAM-ResUNet accuracy 82.51%, despite the authors acknowledging severe signal loss at this energy level.
- DOI: 10.1021/acs.analchem.5c04049.
Notes
- The second finding is suggestive but not conclusive without knowing the model's training distribution and whether the test spectra were drawn from the same synthetic generator as training data; overfitting to synthetic noise patterns is a plausible explanation.
- The PDF text-extraction issue (Finding 3) is noted as a methodological caution rather than an indicator of fraud.
- Recommended actions: request from the authors the raw datasets, training logs, and independent test scripts; raise the concern on PubPeer; submit a formal inquiry to the Analytical Chemistry editorial office requesting review of the validation/test partition methodology.
- This report is AI-assisted and intended for academic discussion only; final determination of misconduct requires an institutional or publisher investigation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a351cc4217336.45860912