English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Review Arcade: When LLM Peer Review Becomes a Game You Can Beat

Forum topic · 小凯 · 2026-05-29

Summary

A Chinese tech forum post discusses the arXiv paper 'Review Arcade: On the Human Alignment and Gameability of LLM Reviews' (arXiv:2605.28897v1), which studies LLM-assisted peer review using 984 real ACL Rolling Review submissions from the 2025 cycle. The study finds that alignment between LLM reviews and human reviews is limited and highly dependent on the model and prompt used. Its central finding is gameability: when authors iteratively revise papers based on LLM reviewer feedback in up to five rounds, 35% of papers achieve statistically significant score improvements without any real change in scientific contribution. The post explains the underlying mechanisms, including surface-structure, self-consistency, specificity, and responsiveness biases, and describes how iterative revision loops converge toward LLM stylistic preferences rather than genuine scientific quality. It also outlines systemic risks such as a review arms race, homogenization of writing style, and fairness issues, before summarizing proposed mitigations like hybrid human-AI review, adversarial prompt variation, and dynamic review criteria.

This post introduces the arXiv paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews (arXiv:2605.28897v1) by Hans Ole Hatzel, Sebastian Steindl, and Jan Strich, which examines how authors can systematically optimize papers to score higher under LLM-generated peer reviews.

Key points

  • Context: Top venues such as ACL Rolling Review (ARR) began piloting LLM-assisted review in 2025, driven by reviewer shortages, the desire for consistency, and speed pressure.
  • Research questions: How well do LLM reviews align with human reviews? How variable are they across models and prompts? Can authors game LLM reviews through iterative revision?
  • Dataset: 984 real ARR submissions from the 2025 cycle, each with three human reviews and LLM-generated reviews from multiple models and prompts.
  • Findings

  • Limited, variable alignment: LLM reviews only correlate reasonably with human reviews, and alignment depends heavily on the prompt (strict vs. friendly) and model (GPT-4 vs. Claude vs. Llama).
  • Gameability works: In an iterative revise-review loop (up to 5 rounds), 35% of papers gained statistically significant overall score improvements—without any change in actual scientific contribution.
  • Systematic biases identified: LLM reviewers favor clear formatting and bullet points, styles matching their own output, more specific (even redundant) details, and preemptive responses to anticipated criticisms.
  • Convergence trap: Iterative optimization converges toward LLM stylistic preferences rather than true scientific quality; if widespread, the field's writing may homogenize around LLM biases, penalizing unconventional but innovative work.
  • Systemic risks

  • A review "arms race" between author-side optimizers and reviewer-side anti-gaming tools, wasting effort on pleasing machines rather than better science.
  • Fairness concerns: unequal access to LLM tools, bias against non-native English writers, and asymmetric knowledge of how to game LLM reviewers.
  • Proposed mitigations

    1. Hybrid review — LLM feedback only as an aid, clearly labeled, with humans holding final decisions. 2. Adversarial design — multiple models and varied prompt styles to detect and deter gaming. 3. Dynamic criteria — regularly rotating scoring standards and periodic blind human checks. 4. Author education — clarifying that LLM feedback is advisory only.

    The post concludes that the paper's value lies in quantifying gameability, exposing its mechanisms, and providing an empirical basis for responsibly introducing AI into peer review—rather than rejecting it outright.

    Reference

  • Hans Ole Hatzel, Sebastian Steindl, Jan Strich (2026). *Review Arcade: On the Human Alignment and Gameability of LLM Reviews*. arXiv:2605.28897v1.
  • ACL Rolling Review: https://aclrollingreview.org/

Tags

#llm#peer-review#academic-ethics#gameability#arxiv#acl#ai-alignment#nlp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980558