Game-the-LLM-Reviewer: Rewriting Papers to Survive AI Peer Review Bias

An open-source skill applies evidence-based rewriting strategies to make papers more appealing to LLM reviewers without altering scientific content, based on analysis of 5 studies involving 120 ICLR 2026 submissions and 4,200 paper versions reviewed by 5 LLMs.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Game-the-LLM-Reviewer: Rewriting Papers to Survive AI Peer Review Bias

Background: AI Reviewers in Peer Review

Researchers now face a new variable: reviewers may feed papers to large language models (LLMs) for evaluation. Studies show that keeping scientific content identical but changing how contributions and results are phrased can alter the scores given by AI reviewers.

Research Foundation: Five Studies on LLM Reviewer Bias

The project synthesizes 5 related studies . One study used 120 anonymous ICLR 2026 submissions to create 4,200 paper versions (including originals) and had 5 different LLMs review them. Across the tested phrasing variations, the wording of experimental results and novelty had the most pronounced effect on scores.

Key Findings

Experimental-result and novelty phrasing influence LLM reviewer scores the most.

The same rewrite does not affect all reviewer models consistently; it can even lower scores for some models.

Therefore the skill applies only relatively stable patterns from the literature—it is not a guaranteed score booster .

Game-the-LLM-Reviewer Skill Design

The skill prioritizes rewriting contribution and result statements ; other strategies are applied on demand. If a meaning-preserving rewrite cannot be found, the original text is kept. The workflow includes:

Evidence check – The agent reads the LaTeX source, cross-references tables and hypotheses, and confirms which claims have supporting evidence.

Targeted rewrite – Small, meaning-preserving edits are made to align with LLM reviewer preferences (e.g., moving key contributions earlier, removing unnecessary hedging).

S6 equivalence check – After editing, the agent verifies that the scientific assessment a human could make remains unchanged. Even if numbers are untouched, a change that strengthens the conclusion is reverted. Consistency between abstract and body is also enforced.

Output – A revised copy is saved alongside changes.md, which documents each rhetorical change, the strategy used, the supporting evidence, and the preserved scientific meaning. If a LaTeX toolchain is present, the agent also checks that the revised manuscript compiles.

Demonstration: Rewriting a Programming Agent Paper Introduction

The authors demonstrate the skill on an introduction for a programming-agent paper. The original text reported that two agents, using different base models, improved their problem-solving rates from 30% to 34% and from 35% to 39% after adding execution memory. The rewrite:

Kept all four numbers exactly the same.

Reordered the narrative: instead of first explaining that agents forget failed repair attempts, it directly introduces execution memory and places the “no model-weight update” fact at the beginning.

Preserved the original background explanation.

The abstract received similar light reordering so that contributions appear earlier, and unnecessary self-deprecating language was made more direct (e.g., “We only tested English queries” → “We tested with English queries. Other languages have not been tested.”). Genuine uncertainty statements were left untouched.

Workflow: Evidence Check, Rewrite, Equivalence Verification

Before editing, the agent reads the main LaTeX file and included section sources to verify evidence for each target statement. After editing, the S6 equivalence check ensures no scientific judgment has been altered. The changes.md log provides a per-edit audit trail.

Installation and Usage

The skill can be added to any compatible coding agent (Claude Code, Codex, etc.). Two installation methods:

# Via repository clone
Clone this repository and install its game-the-llm-reviewer skill:
https://github.com/Michael-Jiahao-Zhang/game-the-llm-reviewer
# Via npx (requires Node.js)
npx skills add Michael-Jiahao-Zhang/game-the-llm-reviewer --skill game-the-llm-reviewer

After installation, invoke the skill on the paper’s entry file (e.g., paper/main.tex). The repository includes a programming-agent introduction example for a quick test.

Philosophy and Availability

The authors explicitly oppose delegating peer-review decisions to LLMs. However, since researchers cannot control whether reviewers use AI assistance, the skill is offered as a defensive measure. The complete skill and accompanying research notes are open-sourced at

https://github.com/Michael-Jiahao-Zhang/game-the-llm-reviewer

.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

academic writingopen-source toolICLR 2026AI reviewer biasequivalence checkingLLM peer reviewpaper rewriting
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.