Summary
This report assesses "FE-YOLOv5: Feature enhancement network based on YOLOv5 for small object detection" by Wang et al. (DOI: 10.1016/j.jvcir.2023.103752, J. Vis. Commun. Image R., 2023). The verdict is "highly suspicious" (academic beautification suspected). The central issue is selective reporting in the abstract: the paper claims YOLOv5 improvements of 2.8% and 2.9% on VisDrone2019 and Tsinghua-Tencent100K respectively, but examination of Tables 1 and 2 shows the 2.8% corresponds to overall AP on VisDrone while the 2.9% actually corresponds to small-object APs on Tsinghua-Tencent (overall AP gain on that dataset is only 1.4%). The two metrics are presented as if they were the same measure. A second concern is that ablation values add perfectly (1.3 + 1.5 = 2.8). All tables report point estimates only, with no standard deviations or significance tests. Image-level analysis was not possible from the text. Confidence is moderate; findings rely on tabular numbers and could be confirmed by verifying the original PDF.
Verdict
Highly suspicious. The paper contains a clear case of selective/misleading reporting in the abstract and presents an implausibly clean ablation pattern. Not proven fabrication, but consistent with academic "beautification."
Key findings
- Selective metric reporting in the abstract: The 2.8% and 2.9% gains attributed to VisDrone2019 and Tsinghua-Tencent100K mix two different metrics (overall AP vs. small-object APs).
- Implausibly perfect ablation additivity: Individual module gains (FEM +1.3, SAM +1.5) sum exactly to the combined gain (2.8) over the YOLOv5 baseline (AP 18.2 → 21.0).
- No uncertainty estimates: Tables 1–4 report only single point estimates; no standard deviation, confidence intervals, or significance tests are provided, despite gains as small as 1.4%.
- Reference list anomalies: Authors' names are missing from the parsed reference entries (e.g., "[1] accurate object detection...", "[2] Computer Vision..."), possibly a PDF parsing artifact but worth verification.
- Image figures not analyzable: Figures 4, 7, and 8 could not be inspected at pixel level from the available text.
Evidence highlights
- Abstract: "Compared to YOLOv5, the … was improved by 2.8% and 2.9%, respectively."
- Table 1 (VisDrone2019): FE-YOLOv5 AP = 21.0, YOLOv5 AP = 18.2 → 21.0 − 18.2 = 2.8% (overall AP).
- Table 2 (Tsinghua-Tencent100K): FE-YOLOv5 AP = 63.6, YOLOv5 AP = 62.2 → 1.4% (overall AP, not 2.9%). APs for FE-YOLOv5 = 52.8, YOLOv5 = 49.9 → 52.8 − 49.9 = 2.9% (small-object APs).
- Table 3 (ablation): Baseline 18.2, +FEM 19.5 (+1.3), +SAM 19.7 (+1.5), +FEM+SAM 21.0 (+2.8). 1.3 + 1.5 = 2.8 exactly.
- DOI: 10.1016/j.jvcir.2023.103752.
Notes
- The VisDrone AP arithmetic confirms the 2.8% claim; the Tsinghua-Tencent claim does not match overall AP and only matches APs.
- Perfect additivity in ablations is statistically uncommon in deep learning but not impossible; it warrants request for raw multi-run logs.
- Authors should be asked for variance estimates on both datasets and clarification of which metric is reported in the abstract.
- Findings are limited to what is visible in the parsed text; original PDF should be inspected for confirmation.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a1e9c37d404e2.82145251