ImpossibleRubrics: LLM-Generated Rubrics Can Be Gamed Up to 64% of the Time

Bowen Qin · hf · 2026-09-16

ImpossibleRubrics: Stress-Testing Generated Rubrics

LLM-generated rubrics are increasingly used as reward signals for rubric-based RL, LLM-as-a-judge evaluation, and automated grading, yet their robustness to adversarial optimization is poorly understood. The authors isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is acknowledging impossibility.

Benchmark design

Key findings

Original post →

More from Models

Models channel →