RobustReview: 1,260-manuscript benchmark exposes rhetorical weakness in AI reviewers, SciCore proposed

Chenguang Wang · hf · 2026-10-02

A new study shows AI reviewers can judge identical science differently depending on wording, rewarding rhetorical optimization over scientific improvement. The authors formalize Rhetorical Robustness as the joint requirement of stability under content-preserving rewrites and discrimination across papers.

Key findings:

The authors then introduce SciCore, a dual-branch reviewer averaging a full-manuscript judgment with one based on an extracted structured "science core." In the primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile while maintaining competitive human alignment.

Original post →

More from Safety

Safety channel →