Study Reveals Rhetoric Can Reward-Hack AI Peer Reviewers, Skewing Scores

UMaryland · hf · 2026-08-14

This paper investigates how rhetorical framing can systematically bias scientific review scores generated by AI. The study reveals that the impact on scores is primarily driven by the AI reviewer's identity, score range, and evaluation strictness, rather than the complexity of rewriting the text.

Original post →

More from Safety

Safety channel →