Wanted to bench Clef as an LLM judge, found a ton of broken rubrics instead

xeophon · x · 2026-10-09

xeophon set out to benchmark Clef as an LLM judge but instead uncovered a large number of broken rubrics, with screenshots attached. A reminder that eval tooling's own rubric quality can't be taken for granted before trusting AI judges.

Original post →

More from Models

Models channel →