Study Reveals Rhetoric Can Reward-Hack AI Peer Reviewers, Skewing Scores
UMaryland · hf · 2026-08-14
This paper investigates how rhetorical framing can systematically bias scientific review scores generated by AI. The study reveals that the impact on scores is primarily driven by the AI reviewer's identity, score range, and evaluation strictness, rather than the complexity of rewriting the text.
More from Safety
- Testing AI to Build an Attack Drone: Safety Guardrails Bypassed — TobyWalsh · 2026-08-14
- LLMs Recognize AI Researchers and Become Less Confident, Study Finds — 机器之心 · 2026-08-14
- Open-Sourcing Agent Skills: Guardrails for Production AI Actions — FunNewspaper5161 · 2026-08-14
- Over 800 Fake AI Skills and MCP Servers Found Delivering Malware — HaktanSuren · 2026-08-14
- Beware: Malicious Google Ads Mimic ChatGPT to Phish Users via Windows Run — CCB0x45 · 2026-08-14
- Cooperative AI Seminar: Solving AI Game Theory Dilemmas with Safe Pareto Improvements — xuanalogue · 2026-08-14