Game-the-LLM-Reviewer: Rewriting Papers to Survive AI Peer Review Bias
An open-source skill applies evidence-based rewriting strategies to make papers more appealing to LLM reviewers without altering scientific content, based on analysis of 5 studies involving 120 ICLR 2026 submissions and 4,200 paper versions reviewed by 5 LLMs.
Background: AI Reviewers in Peer Review
Researchers now face a new variable: reviewers may feed papers to large language models (LLMs) for evaluation. Studies show that keeping scientific content identical but changing how contributions and results are phrased can alter the scores given by AI reviewers.
Research Foundation: Five Studies on LLM Reviewer Bias
The project synthesizes 5 related studies . One study used 120 anonymous ICLR 2026 submissions to create 4,200 paper versions (including originals) and had 5 different LLMs review them. Across the tested phrasing variations, the wording of experimental results and novelty had the most pronounced effect on scores.
Key Findings
Experimental-result and novelty phrasing influence LLM reviewer scores the most.
The same rewrite does not affect all reviewer models consistently; it can even lower scores for some models.
Therefore the skill applies only relatively stable patterns from the literature—it is not a guaranteed score booster .
Game-the-LLM-Reviewer Skill Design
The skill prioritizes rewriting contribution and result statements ; other strategies are applied on demand. If a meaning-preserving rewrite cannot be found, the original text is kept. The workflow includes:
Evidence check – The agent reads the LaTeX source, cross-references tables and hypotheses, and confirms which claims have supporting evidence.
Targeted rewrite – Small, meaning-preserving edits are made to align with LLM reviewer preferences (e.g., moving key contributions earlier, removing unnecessary hedging).
S6 equivalence check – After editing, the agent verifies that the scientific assessment a human could make remains unchanged. Even if numbers are untouched, a change that strengthens the conclusion is reverted. Consistency between abstract and body is also enforced.
Output – A revised copy is saved alongside changes.md, which documents each rhetorical change, the strategy used, the supporting evidence, and the preserved scientific meaning. If a LaTeX toolchain is present, the agent also checks that the revised manuscript compiles.
Demonstration: Rewriting a Programming Agent Paper Introduction
The authors demonstrate the skill on an introduction for a programming-agent paper. The original text reported that two agents, using different base models, improved their problem-solving rates from 30% to 34% and from 35% to 39% after adding execution memory. The rewrite:
Kept all four numbers exactly the same.
Reordered the narrative: instead of first explaining that agents forget failed repair attempts, it directly introduces execution memory and places the “no model-weight update” fact at the beginning.
Preserved the original background explanation.
The abstract received similar light reordering so that contributions appear earlier, and unnecessary self-deprecating language was made more direct (e.g., “We only tested English queries” → “We tested with English queries. Other languages have not been tested.”). Genuine uncertainty statements were left untouched.
Workflow: Evidence Check, Rewrite, Equivalence Verification
Before editing, the agent reads the main LaTeX file and included section sources to verify evidence for each target statement. After editing, the S6 equivalence check ensures no scientific judgment has been altered. The changes.md log provides a per-edit audit trail.
Installation and Usage
The skill can be added to any compatible coding agent (Claude Code, Codex, etc.). Two installation methods:
# Via repository clone
Clone this repository and install its game-the-llm-reviewer skill:
https://github.com/Michael-Jiahao-Zhang/game-the-llm-reviewer # Via npx (requires Node.js)
npx skills add Michael-Jiahao-Zhang/game-the-llm-reviewer --skill game-the-llm-reviewerAfter installation, invoke the skill on the paper’s entry file (e.g., paper/main.tex). The repository includes a programming-agent introduction example for a quick test.
Philosophy and Availability
The authors explicitly oppose delegating peer-review decisions to LLMs. However, since researchers cannot control whether reviewers use AI assistance, the skill is offered as a defensive measure. The complete skill and accompanying research notes are open-sourced at
https://github.com/Michael-Jiahao-Zhang/game-the-llm-reviewer.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
