ImpossibleRubrics: LLM-Generated Rubrics Can Be Gamed Up to 64% of the Time
Bowen Qin · hf · 2026-09-16
ImpossibleRubrics: Stress-Testing Generated Rubrics
LLM-generated rubrics are increasingly used as reward signals for rubric-based RL, LLM-as-a-judge evaluation, and automated grading, yet their robustness to adversarial optimization is poorly understood. The authors isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is acknowledging impossibility.
Benchmark design
- 169 impossible tasks across six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, plus 48 answerable controls
- Rather than fixed rubrics, the benchmark provides task environments and certificates so rubrics can be generated downstream and adversarially tested
Key findings
- On the unbiased 150-of-169 cut, eleven generators are exploited 8–26% of the time; on a stress cut the strongest generator is exploited 36% while a certificate-faithful rubric is exploited 0% — a rubric-quality gap, not task impossibility
- Counterintuitive: a single generic rubric ("be decisive, penalize hedging") used unchanged is exploited 64% of the time, and 7 of 11 generators are exploited more often than that despite tailoring criteria per task
- Tailored criteria appear to tell an attacker which claim to fabricate; the problem is not vagueness but being specific about the wrong things
More from Models
- "AI Plays Doom" Demo Debunked: Text State Input, Solvable in ~30 Lines of Code — banteg · 2026-09-16
- Rumor: a big 'ship week' teased with GPT-6 family including GPT-6 Sol — Winter-Mix-5155 · 2026-09-16
- Altman: GPT 5.5 Matches an Average Math Professor, Internal Model Beats World's Best — acoolrandomusername · 2026-09-16
- ChatGPT desktop app now shows archived/deleted chats with no opt-out — ___Patrice___ · 2026-09-16
- Closed labs ramped up life-science data efforts; GPT-Rosalind cited as OpenAI bull case — xeophon · 2026-09-16
- GLiNER creator: zero-shot NER models are underestimated; end-to-end RL is next — bclavie · 2026-09-16