English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Geng Report 6a263faf: Suspicious Findings in "The Asymmetric Vulnerability: Bypassing LLM Defenses via Guardrail-Model Mismatch"

Academic fraud report · Geng Detector

Summary

This report assesses the WWW '26 paper by Junyi Wang, Zhibin Zhu, and Chuanyi Liu (Harbin Institute of Technology, Shenzhen & Peng Cheng Laboratory; DOI: 10.1145/3774904.3792438) and assigns a verdict of "Suspicious" rather than outright fraud. The main tables (Tables 1 and 2) demonstrate unusually clean arithmetic with ASR values such as 87.5%, 81.5%, 90.5%, 92.5%, 98.5%, and 99.5% — all exact multiples of 0.5% on a 200-sample base — suggesting real experimental runs rather than fabricated data. However, Appendix A (Table 3) reports integer percentages on a sample of n=553, which is statistically inconsistent without rounding disclosure. More critically, Appendix B explicitly states that "each configuration was executed twice, and the run in which the attack succeeded was selected," a textbook cherry-picking design that inflates ASR and undermines Figures 3 and 7. Additionally, GPT-4o serves as both target model and sole evaluator, creating a self-judging loop. Evidence is methodological rather than fabrication-based; conclusions about the RepMism framework's robustness remain unverified pending raw multi-run logs.

Verdict

Suspicious (🟡). The paper does not exhibit hallmarks of systematic data fabrication. Instead, it contains serious methodological flaws — most notably an explicit cherry-picking protocol and a self-referential evaluation loop — that meaningfully inflate reported attack success rates and warrant caution before accepting the RepMism framework's claimed robustness.

Key findings

  • Clean arithmetic on n=200 (Tables 1, 2): Every reported ASR is a multiple of 0.5% (87.5%, 81.5%, 90.5%, 92.5%, 98.5%, 99.5%). This corresponds to exact integer counts (175, 163, 181, 185, 197, 199 / 200). Such perfect consistency argues *against* fabrication and *for* genuine sampling.
  • Integer percentages on n=553 (Table 3, Appendix A): Values such as 10%, 32%, 58%, 95%, 99% cannot arise from a 553-sample run without undisclosed rounding (e.g., 10% implies 55.3 cases). Without per-cell counts, the confusion matrix is not reproducible.
  • Cherry-picking protocol (Appendix B): "Each configuration was executed twice, and the run in which the attack succeeded was selected." This is a textbook lottery-style selection bias. Two runs do not characterize stochasticity of closed-source LLMs; the reported ASR upper-bounds rather than estimates true success probability. Core Figures 3 and 7 are therefore questionable.
  • Self-referential evaluation (Section 4.3, Appendix D): GPT-4o is both target and sole judge. The reported Pearson correlation of r = 0.989 with p = 3.52 × 10⁻¹⁶⁸ is plausible at high correlation but does not eliminate systematic LLM-judge bias.
  • Timeline consistency: Model references (GPT-4o, GPT-o3, DeepSeek-r1, Gemini-2.5, Qwen3-8B) and cited arXiv preprints (e.g., 2508.14070, 2025) are temporally consistent with a 2026 WWW submission. No anachronisms detected.
  • Evidence highlights

  • n=200 / 0.5%-granularity cross-check:
  • 87.5% → 175/200
  • 81.5% → 163/200
  • 90.5% → 181/200
  • 92.5% → 185/200
  • 98.5% → 197/200
  • 99.5% → 199/200
  • n=553 / integer rounding inconsistency: 10% on 553 samples requires 55.3 cases; nearest integers lie in [52, 57].
  • Pearson r = 0.989, p = 3.52 × 10⁻¹⁶⁸ (GPT-4o vs. human annotators, Appendix D).
  • DOI: 10.1145/3774904.3792438.

Notes

The most actionable follow-up is to request the authors' complete multi-run logs (not just the two-run "winning" subsets) for Figure 3's perturbation-strength curve and Figure 7. If the underlying distribution of ASR across runs is bimodal or highly skewed, the cherry-picking claim would be substantiated. Conversely, if independent re-runs reproduce the curve within tolerance, the methodological criticism weakens. Until such logs are provided, the paper's quantitative claims should be treated as upper bounds. No image-based analysis was possible (no figures supplied for pixel inspection).

Tags

#academic-fraud#methodology#cherry-picking#statistics#llm-security#peer-review#evaluation-bias#data-integrity

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a263faf8e2be1.28776446