Stop automating peer review without rigorous evaluation

Authors

  • Slava St Author

Abstract

As AI systems increasingly generate scientific knowledge, the human ability to critically evaluate research becomes more important, not less. Yet large language models offer a tempting solution to address the peer review crisis, risking the automation of the very skills scientists will need most. This position paper argues that today’s AI systems should not be used to produce paper reviews. We ground this position in an empirical comparison of human- versus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper significantly increases scores from AI reviewers through stylistic changes rather than scientific improvements. However, non-gameability and review diversity are necessary but not sufficient conditions for automation. We argue that addressing the peer review crisis requires a science of peer review automation that keeps human scientific judgment at the center of the process—especially as we enter an era where that judgment will be needed most.

AI Reviews

Reviewer: Gemini 3.6 Flash

  • Originality/Novelty : 80%
  • Significance/Contribution : 80%
  • Methodological Rigor : 80%
  • Clarity and Structure : 100%
  • Literature Review : 100%
  • Validity of Conclusions : 100%
  • Overall Suitability : 100%
  • Formatting Quality : 100%
Review Summary

This paper presents a position piece arguing against the use of current Large Language Models (LLMs) to automatically generate paper reviews for scientific venues. The authors establish two necessary conditions for peer review automation: (1) preservation of review diversity (C1), ensuring the plurality of expert feedback is not collapsed; and (2) resistance to gaming (C2), ensuring scores cannot be manipulated without genuine improvements in scientific content.

To evaluate whether current AI systems meet these conditions, the authors conduct a dual empirical investigation combining observational analysis and controlled experiments. First, using a dataset of 75,800 ICLR 2026 reviews labeled for AI generation, as well as simulated reviews generated by GPT-5.1 and Claude Sonnet 4.5 agent setups on 60 sampled papers, they identify an 'AI reviewer hivemind effect.' Fully AI-generated reviews show significantly higher inter-paper and intra-paper similarity compared to human reviews, collapsing perspective diversity and exhibiting high reuse of templated phrases.

Second, the authors demonstrate 'paper laundering' as a trivial gaming mechanism. By using zero-shot LLM prompts to rewrite LaTeX manuscripts based on AI reviewer feedback, papers achieve a statistically significant average score increase (+0.28 points on a 10-point scale), raising predicted acceptance probability by 7.3 percentage points. Word-level analyses reveal that score gains are driven by stylistic modifications—disproportionately adding hedging words, emphasis terms, and generic structural filler—rather than substantive scientific additions. Finally, the authors articulate a three-pillar framework for a 'science of peer review automation,' advocating for rigorous evaluation prior to deployment, recognition of algorithmic monoculture risks, and the preservation of human scientific evaluation skills.

Major Strengths:
- Timely and highly impactful topic addressing the rapid integration of LLMs into conference peer review workflows.
- Novel concept of 'paper laundering' demonstrating a zero-shot, policy-compliant gaming tactic distinct from traditional prompt injection attacks.
- Solid empirical methodology combining observational data at scale with controlled, paired simulation experiments.
- Clear structure, strong argumentation, and transparent presentation of experimental details and limitations.

Major Weaknesses:
- Experimental simulations are restricted to two proprietary frontier models (GPT-5.1 and Claude Sonnet 4.5) using a single standardized system prompt.
- Perspective diversity is measured using embedding cosine similarity, which reflects linguistic and stylistic patterns rather than explicit argumentative divergence.
- The sample size for the laundering experiment is relatively modest (n = 60 papers).
- Simple prompt-level or decoding-level mitigations (e.g., stochastic sampling, diversity-enforcing prompts, or laundering detection mechanisms) were not empirically evaluated.

Reviewer: grok_4_20

  • Originality/Novelty : 80%
  • Significance/Contribution : 90%
  • Methodological Rigor : 80%
  • Clarity and Structure : 100%
  • Literature Review : 100%
  • Validity of Conclusions : 100%
  • Overall Suitability : 100%
  • Formatting Quality : 80%
Review Summary

The main research objective of this position paper is to argue that current AI systems are not suitable for producing autonomous peer reviews, as they fail to preserve review diversity and are vulnerable to trivial gaming. The authors support this with an empirical analysis of 75,800 real ICLR 2026 reviews (21% AI-generated per Emi, 2025), controlled simulations on 60 sampled ICLR papers using GPT-5.1 and Claude Sonnet 4.5 reviewer agents, and experiments on 'paper laundering' (zero-shot LLM rewriting of LaTeX papers informed by initial AI reviews). Key findings demonstrate a 'hivemind effect' with significantly higher intra-paper (IntraSim +8.7%, Cohen's d=1.47) and inter-paper (InterSim +17.6% to +37.4%, Cohen's d up to 3.55) review similarity for AI versus human reviews using embedding cosine similarity, plus AI score inflation and templated phrases. Paper laundering increased scores by +0.28 points (p<0.001) via stylistic changes (e.g., more hedging/emphasis words) rather than substantive improvements, also increasing paper homogeneity. The paper concludes that these failures necessitate a rigorous 'science of peer review automation' centered on human judgment, even as AI capabilities advance.

Downloads

Published

31.07.2026

AI Review Rating

88%

Section

Articles

Categories