English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Academic Fraud Detection Report: NoteIt (UIST '25) — Inconsistent Metrics and Suspicious Study Results

Academic fraud report · Geng Detector

Summary

This report evaluates the paper "NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding" (DOI: 10.1145/3746059.3747626, UIST '25), authored by Running Zhao et al. Verdict: Highly Suspicious. The primary finding is a verified mathematical inconsistency in Table 1: the reported F1 scores do not match the reported Precision/Recall values under the standard F1 formula. For the dynamic scenes (Camera perspective manipulation), the paper reports Recall=67.94%, Precision=90.96%, F1=70.86%, but the mathematically correct F1 is 77.78%. For static scenes, the reported F1=91.63% is inconsistent with the back-calculated 92.04% given Recall=91.88% and Precision=92.20%. These discrepancies indicate the figures were not derived from a genuine confusion matrix. A secondary finding (unconfirmed) flags the unusual perfection of the user study in Table 2, where all p-values are below 0.001 with n=36, and standard errors are suspiciously uniform. Citation clustering concerns remain speculative due to incomplete reference data. Confidence in Finding 1 is high; confidence in Findings 2–3 is limited pending access to raw data.

Verdict

Highly Suspicious (🟠). The paper contains a mathematically verified inconsistency in its core evaluation metrics (Table 1), undermining the credibility of the reported quantitative results. Additional concerns regarding user-study statistics and citation patterns are flagged but require corroboration with raw data and a full reference list. Final adjudication requires institutional investigation.

Key findings

  • Verified mathematical inconsistency in Table 1 (F1 vs. Precision/Recall). The reported F1 scores cannot be reproduced from the reported Precision and Recall using the standard harmonic-mean formula, suggesting fabricated or post-hoc adjusted figures.
  • Suspiciously uniform user-study results in Table 2 (unconfirmed). All pairwise comparisons yield p < 0.001 with n = 36, and standard errors cluster within a narrow range (0.07–0.17). This pattern is atypical for Likert-style subjective ratings and may indicate data smoothing or selective reporting.
  • Possible citation clustering (unconfirmed). The full reference list was not available to verify suspected self-referential citation practices by the same author group within a narrow publication window.
  • Evidence highlights

  • Table 1 — Dynamic scene (Camera perspective manipulation): Reported Recall = 67.94%, Precision = 90.96%, F1 = 70.86%. Back-calculated F1 = 2 × (0.9096 × 0.6794) / (0.9096 + 0.6794) = 77.78%. Discrepancy ≈ 6.92 percentage points.
  • Table 1 — Static scene: Reported Recall = 91.88%, Precision = 92.20%, F1 = 91.63%. Back-calculated F1 = 2 × (0.9220 × 0.9188) / (0.9220 + 0.9188) = 92.04%. Discrepancy ≈ 0.41 percentage points.
  • Table 2 — Pairwise comparisons: 7 bidirectional comparisons, all p < 0.001; sample size n = 36; non-parametric test referenced. NoteIt SE values reportedly span only 0.07–0.17.
  • DOI: 10.1145/3746059.3747626
  • Venue: The 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), 2025.
  • Notes

  • Confidence: High for Finding 1 (closed-form arithmetic check). Low for Findings 2 and 3 (require raw survey responses and complete bibliography for confirmation).
  • Limitations: The arithmetic error in Finding 1 could in principle stem from a typographical/typesetting error or from a non-standard F1 definition, though the latter is not indicated in the paper. The user-study concern cannot be resolved without access to the raw Likert responses and the exact statistical test used.
  • Recommended actions: Request raw data, calculation spreadsheets, and survey responses from the authors; raise concerns via PubPeer; report to the ACM UIST 2025 program committee for editorial review.
  • Disclaimer: This report is AI-assisted and intended for academic discussion only. Final determination of misconduct requires institutional investigation.

Tags

#academic-fraud#data-inconsistency#f1-mismatch#uist-2025#hci#multimodal#statistics#user-study

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a393f866e6b70.38473524