This post introduces the arXiv paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews (arXiv:2605.28897v1) by Hans Ole Hatzel, Sebastian Steindl, and Jan Strich, which examines how authors can systematically optimize papers to score higher under LLM-generated peer reviews.
Key points
- Context: Top venues such as ACL Rolling Review (ARR) began piloting LLM-assisted review in 2025, driven by reviewer shortages, the desire for consistency, and speed pressure.
- Research questions: How well do LLM reviews align with human reviews? How variable are they across models and prompts? Can authors game LLM reviews through iterative revision?
- Dataset: 984 real ARR submissions from the 2025 cycle, each with three human reviews and LLM-generated reviews from multiple models and prompts.
- Limited, variable alignment: LLM reviews only correlate reasonably with human reviews, and alignment depends heavily on the prompt (strict vs. friendly) and model (GPT-4 vs. Claude vs. Llama).
- Gameability works: In an iterative revise-review loop (up to 5 rounds), 35% of papers gained statistically significant overall score improvements—without any change in actual scientific contribution.
- Systematic biases identified: LLM reviewers favor clear formatting and bullet points, styles matching their own output, more specific (even redundant) details, and preemptive responses to anticipated criticisms.
- Convergence trap: Iterative optimization converges toward LLM stylistic preferences rather than true scientific quality; if widespread, the field's writing may homogenize around LLM biases, penalizing unconventional but innovative work.
- A review "arms race" between author-side optimizers and reviewer-side anti-gaming tools, wasting effort on pleasing machines rather than better science.
- Fairness concerns: unequal access to LLM tools, bias against non-native English writers, and asymmetric knowledge of how to game LLM reviewers.
- Hans Ole Hatzel, Sebastian Steindl, Jan Strich (2026). *Review Arcade: On the Human Alignment and Gameability of LLM Reviews*. arXiv:2605.28897v1.
- ACL Rolling Review: https://aclrollingreview.org/
Findings
Systemic risks
Proposed mitigations
1. Hybrid review — LLM feedback only as an aid, clearly labeled, with humans holding final decisions. 2. Adversarial design — multiple models and varied prompt styles to detect and deter gaming. 3. Dynamic criteria — regularly rotating scoring standards and periodic blind human checks. 4. Author education — clarifying that LLM feedback is advisory only.
The post concludes that the paper's value lies in quantifying gameability, exposing its mechanisms, and providing an empirical basis for responsibly introducing AI into peer review—rather than rejecting it outright.