RobustReview: 1,260-manuscript benchmark exposes rhetorical weakness in AI reviewers, SciCore proposed
Chenguang Wang · hf · 2026-10-02
A new study shows AI reviewers can judge identical science differently depending on wording, rewarding rhetorical optimization over scientific improvement. The authors formalize Rhetorical Robustness as the joint requirement of stability under content-preserving rewrites and discrimination across papers.
Key findings:
- RobustReview benchmark: 1,260 controlled manuscript versions, 30 reviewer configurations evaluated.
- Uncovers "false robustness": low rewrite sensitivity can coexist with score collapse across papers.
- Human alignment and rhetorical robustness rank reviewers differently; content-focused prompting doesn't consistently help.
The authors then introduce SciCore, a dual-branch reviewer averaging a full-manuscript judgment with one based on an extracted structured "science core." In the primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile while maintaining competitive human alignment.
More from Safety
- OpenRouter launches Security Center after finding 1,000+ dormant API keys across 85 employees — AccBalanced · 2026-10-02
- Open weights called "dangerous" — echoing every tech that shifted information power — 0xAllen_ · 2026-10-02
- Abliterated Large V2 lands on Venice: refusal-free AI for red teamers, anonymously — 0xAllen_ · 2026-10-02
- Microsoft's 2026 Digital Defense Report: AI accelerates attacks, 52.2% of intrusions pursue credential theft — AccBalanced · 2026-10-02
- Meta's Muse agent joins your Tailscale tailnet as a node, with existing access controls intact — AccBalanced · 2026-10-02
- Zero-privilege default architecture is the #1 way to limit prompt injection blast radius — AccBalanced · 2026-10-02