Verdict
Highly Suspicious (🟠). The paper contains a mathematically verified inconsistency in its core evaluation metrics (Table 1), undermining the credibility of the reported quantitative results. Additional concerns regarding user-study statistics and citation patterns are flagged but require corroboration with raw data and a full reference list. Final adjudication requires institutional investigation.
Key findings
- Verified mathematical inconsistency in Table 1 (F1 vs. Precision/Recall). The reported F1 scores cannot be reproduced from the reported Precision and Recall using the standard harmonic-mean formula, suggesting fabricated or post-hoc adjusted figures.
- Suspiciously uniform user-study results in Table 2 (unconfirmed). All pairwise comparisons yield p < 0.001 with n = 36, and standard errors cluster within a narrow range (0.07–0.17). This pattern is atypical for Likert-style subjective ratings and may indicate data smoothing or selective reporting.
- Possible citation clustering (unconfirmed). The full reference list was not available to verify suspected self-referential citation practices by the same author group within a narrow publication window.
- Table 1 — Dynamic scene (Camera perspective manipulation): Reported Recall = 67.94%, Precision = 90.96%, F1 = 70.86%. Back-calculated F1 = 2 × (0.9096 × 0.6794) / (0.9096 + 0.6794) = 77.78%. Discrepancy ≈ 6.92 percentage points.
- Table 1 — Static scene: Reported Recall = 91.88%, Precision = 92.20%, F1 = 91.63%. Back-calculated F1 = 2 × (0.9220 × 0.9188) / (0.9220 + 0.9188) = 92.04%. Discrepancy ≈ 0.41 percentage points.
- Table 2 — Pairwise comparisons: 7 bidirectional comparisons, all p < 0.001; sample size n = 36; non-parametric test referenced. NoteIt SE values reportedly span only 0.07–0.17.
- DOI: 10.1145/3746059.3747626
- Venue: The 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), 2025.
- Confidence: High for Finding 1 (closed-form arithmetic check). Low for Findings 2 and 3 (require raw survey responses and complete bibliography for confirmation).
- Limitations: The arithmetic error in Finding 1 could in principle stem from a typographical/typesetting error or from a non-standard F1 definition, though the latter is not indicated in the paper. The user-study concern cannot be resolved without access to the raw Likert responses and the exact statistical test used.
- Recommended actions: Request raw data, calculation spreadsheets, and survey responses from the authors; raise concerns via PubPeer; report to the ACM UIST 2025 program committee for editorial review.
- Disclaimer: This report is AI-assisted and intended for academic discussion only. Final determination of misconduct requires institutional investigation.