English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Integrity Concerns in “Replication and Exploration of Generative Retrieval over Dynamic Corpora” (DOI: 10.1145/3726302.3730314)

Academic fraud report · Geng Detector

Summary

This report identifies substantial internal inconsistencies in the reported evaluation of “Replication and Exploration of Generative Retrieval over Dynamic Corpora” (DOI: 10.1145/3726302.3730314), leading to an overall assessment of highly suspicious rather than a definitive finding of research misconduct. The clearest issue is a numerical contradiction: the text claims that MDGR is 3.6× faster than SEAL, whereas the cited latency values imply approximately 2.57× at Tok-K = 10 and 2.43× at Tok-K = 100. A second issue concerns inconsistent NQ baseline results: SEAL (N-gram) is reported as Hit@10 = 0.809 on D0 in Table 1, while N-gram is reported as 0.753 on D0 in the Table 4 ablation. The paper also states that all table results are significant at p < 0.05 but reports no p-values, test statistics, or variability estimates, so that claim cannot be independently checked. Table 2 additionally includes a D0 column despite being described as covering newly added document sets D1 to D5. Together, these issues indicate deficient numerical checking, table assembly, or experimental documentation. They warrant clarification and access to raw logs, but the available text does not establish intentional fabrication.

Verdict

Overall assessment: highly suspicious, but not a definitive determination of academic misconduct. The paper contains multiple reproducible internal inconsistencies affecting its central performance claim, baseline comparability, and statistical reporting. The evidence supports serious concerns about numerical checking, data handling, or manuscript assembly. It does not, by itself, establish intentional fabrication. A formal conclusion would require the authors’ raw logs, source tables, experimental configuration, and statistical-analysis files.

Key findings

  • Unsupported speed-up claim: Section 6.3 and Table 7 state that MDGR is 3.6× faster than SEAL, but the reported latency values do not yield that ratio. At Tok-K = 10, 619 ms divided by 241 ms is approximately 2.57×; at Tok-K = 100, 778 ms divided by 320 ms is approximately 2.43×.
  • Inconsistent NQ baseline: Table 1 reports SEAL (N-gram) Hit@10 = 0.809 on D0, whereas Table 4 reports N-gram D0 Hit@10 = 0.753. Because the dataset, metric, and model family are identified as the same, this 0.056 difference requires explanation.
  • Statistical claim is not auditable: Section 3.4 claims that all experimental results in the tables are significant at p < 0.05, but the tables provide no p-values, t-values, F-values, standard deviations, or standard errors. The basis of the significance claim therefore cannot be verified.
  • Questionable table organization: Table 2 is described as evaluating newly added document sets D1 to D5 but includes a D0 column, with values such as BM25 = 0.647 that match the earlier D0 results.
  • Evidence highlights

  • Reported claim: “MDGR retrieves documents 3.6× faster than SEAL.”
  • Table 7 at Tok-K = 10: SEAL = 619 ms; MDGR = 241 ms; implied speed-up = 619 / 241 ≈ 2.57×.
  • Table 7 at Tok-K = 100: SEAL = 778 ms; MDGR = 320 ms; implied speed-up = 778 / 320 ≈ 2.43×.
  • Table 1, NQ, D0: SEAL (N-gram) Hit@10 = 0.809.
  • Table 4, NQ, D0: N-gram Hit@10 = 0.753.
  • Absolute baseline discrepancy: 0.809 − 0.753 = 0.056.
  • Reported threshold: p < 0.05; no corresponding inferential statistics are supplied in the cited tables.

Notes

The speed-up calculations above assume that lower latency alone defines “faster”; the paper’s wording and the Table 7 values are consistent with using the SEAL-to-MDGR latency ratio. Variability, hardware conditions, batching, warm-up, confidence intervals, and repeated-run methodology are not provided, so the reliability of the latency measurements cannot be evaluated.

Several explanations remain possible, including transcription errors, different experimental settings, accidental inclusion of unrelated runs, table-copying errors, or unsupported emphasis in the prose. These possibilities prevent a categorical fraud finding from the text alone. Appropriate follow-up requests include the raw SEAL and MDGR timing logs, complete source data for Tables 1, 2, 4, and 7, clarification of whether the reported rows use identical settings, and the statistical procedure and outputs supporting the p < 0.05 statement. The limitations of this text-only analysis should also be acknowledged: it evaluates internal numerical and reporting consistency, not the provenance of every underlying experiment.

Tags

#academic-fraud#data-inconsistency#baseline-discrepancy#statistical-reporting#unsupported-claim#table-quality#retrieval-systems

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a32c2a6c2c342.99315834