English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Review Arcade: Human Alignment and Gaming of LLM Peer Reviews

Forum topic · 小凯 · 2026-05-30

Summary

This arXiv paper (2605.28897) by Hans Ole Hatzel, Sebastian Steindl, and Jan Strich empirically evaluates LLM-generated peer reviews from both the author and reviewer perspectives, using papers from the 2025 ACL Rolling Review (ARR). The study finds limited alignment between LLM reviews and human reviews: while alignment can be reasonable in the best case, it varies substantially across prompts and models. The authors also investigate a scenario where authors use an iterative draft-revise workflow to improve their submissions based on LLM feedback. This 'gaming' of LLM reviews proves effective in specific scenarios, producing a statistically significant increase in overall review scores for up to 35% of papers. The findings raise concerns about deploying LLM reviews in academic peer review, since both reviewers and authors may use LLM assistance, potentially distorting evaluation outcomes.

Paper Overview

Research Area: LLM Evaluation Authors: Hans Ole Hatzel, Sebastian Steindl, Jan Strich Published: 2026-05-30 arXiv: 2605.28897

Abstract

LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers are using LLM-assistance, but also that authors use LLMs to revise their papers before submitting. In this work, we perform empirical experiments on papers from the 2025 ACL Rolling Review (ARR) to evaluate LLM reviews from both the author and the reviewer perspective.

First, we identify a limited alignment of LLM reviews with human ones. In the best-case scenario, the alignment is reasonable. However, we also find that LLM-human alignment varies substantially across prompts and models.

Finally, we investigate the scenario in which the author uses an iterative draft-revise workflow to improve the submission according to the LLM review. We find that this "gaming" of LLM reviews can be effective in specific scenarios, leading to a statistically significant increase of overall scores for up to 35% of papers.

Key Takeaways

  • LLM reviews show only limited alignment with human peer reviews; best-case alignment is reasonable but inconsistent.
  • Alignment quality varies significantly depending on the prompt and model used.
  • Authors can exploit LLM reviews through iterative draft-revise workflows, significantly boosting overall scores for up to 35% of papers.
  • These results caution against naive deployment of LLM-based reviewing in academic conferences, given that both sides of the review process may leverage LLMs.

Tags

#llm#peer-review#arxiv#evaluation#academic-publishing#acl-rolling-review#gaming

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980560