Summary
This report flags the 2017 ACM Multimedia workshop paper "Multispectral Object Detection for Autonomous Vehicles" by Karasawa, Watanabe, Ha, Tejero-De-Pablos, Ushiku, and Harada as highly suspect due to a seriously flawed experimental comparison, though without evidence of numerical fabrication. Arithmetic checks of mean Average Precision (mAP) values in Tables 2, 3, 5, and 6 match the reported per-class APs (e.g., Table 2 RGB: 0.656 ≈ 0.66; NIR: 0.616 ≈ 0.62; Table 3 Ensemble: 0.786 ≈ 0.79), so quantitative reporting appears internally consistent. The core concern is asymmetric dataset allocation in Section 3 and Section 5.2: the six-channel baseline is trained on only ~1,553 deliberately selected misaligned images, while the proposed Ensemble method's sub-models are trained on the full ~7,512 original RGB/FIR/MIR/NIR images. This 'horse-racing' comparison likely inflates the reported 13% Ensemble advantage by conflating data quantity/quality effects with architectural merit. The authors further attribute the baseline's failure to feature-fusion difficulty, an interpretation unsupported given the training-set disparity. No timestamp anomalies were found in the listed imaging hardware. Limitations: findings rely on textual descriptions; underlying code or raw splits were not inspected.
Verdict
🟠
Highly suspect — no evidence of numerical data fabrication, but the central comparative claim rests on an evidently unfair experimental design that confounds data quantity/quality with method choice.
Key findings
- Arithmetic consistency (no fabrication detected): Per-class APs in Tables 2, 3, 5, and 6 average to the reported mAP values within rounding tolerance. Example: Table 2 RGB row — (0.55 + 0.77 + 0.64 + 0.58 + 0.74) / 5 = 0.656, reported as 0.66. Table 2 NIR — 3.08 / 5 = 0.616, reported as 0.62. Table 3 Ensemble — 3.93 / 5 = 0.786, reported as 0.79.
- Asymmetric training data between baseline and proposed method (core red flag): Section 3 states that the authors selected 1,446 well-aligned images as test images and 1,553 poorly-aligned images as training images for the six-channel experiment, whereas the Ensemble sub-models are trained "separately using the original (not merged) RGB, FIR, MIR, and NIR images" — i.e., on the full ~7,512-image dataset.
- Confounded performance comparison: The reported ~13% mAP advantage of Ensemble over the six-channel baseline cannot be cleanly attributed to the ensemble architecture, because the baseline was starved of training data and given only low-quality (misaligned) samples, while the proposed method consumed the full corpus.
- Unsupported theoretical attribution: In Section 5.2.1, the authors attribute the baseline's failure to "not able to train a meaningful feature extraction model" / feature-fusion difficulty. Given the data disparity, this interpretation is not warranted by the experiments as conducted.
- Hardware timeline is plausible: RGB Logicool C920R (released 2012), FIR Nippon Avionics InfReC R500, NIR Xenics Xeva-1.7-320 — all predate the 2017 publication; no anachronistic devices were found.
Evidence highlights
- DOI: 10.1145/3126686.3126727
- Numerical check, Table 2 RGB: (0.55 + 0.77 + 0.64 + 0.58 + 0.74) / 5 = 0.656 → reported 0.66 ✓
- Numerical check, Table 2 NIR: 3.08 / 5 = 0.616 → reported 0.62 ✓
- Numerical check, Table 3 Ensemble: 3.93 / 5 = 0.786 → reported 0.79 ✓
- Key quotation (Section 3): "we visually reviewed … and chose 1,446 images whose alignments are correct … as test images and 1,553 images whose alignments are incorrect … as training images for the experiment using six channels."
- Key contradiction (Section 4.1 / 5.2.1): Ensemble sub-models "trained separately using the original (not merged) RGB, FIR, MIR, and NIR images" vs. six-channel baseline trained only on the 1,553 misaligned subset.
- Reported gap: ~13% mAP in favor of Ensemble, used to justify the proposed method over six-channel input.
Notes
- This report identifies methodological/design concerns, not proven misconduct. Final determination of academic misconduct requires investigation by an appropriate body (e.g., the publisher or the authors' institution).
- Recommended follow-ups: (1) request the authors to provide an apples-to-apples comparison, e.g., training both Ensemble sub-models and the six-channel model on identical data partitions (either the full set or the 1,553-image subset); (2) raise the asymmetry concern on PubPeer or directly with the ACM workshop chairs; (3) request release of the exact dataset splits and training scripts to enable independent verification.
- Limitations: analysis is based on the published PDF; raw training logs, code, and dataset files were not inspected. Any conclusion about intent (deliberate bias vs. oversight) cannot be drawn from the text alone.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a1d3feb3dc326.91522171