Wanted to bench Clef as an LLM judge, found a ton of broken rubrics instead
xeophon · x · 2026-10-09
xeophon set out to benchmark Clef as an LLM judge but instead uncovered a large number of broken rubrics, with screenshots attached. A reminder that eval tooling's own rubric quality can't be taken for granted before trusting AI judges.
More from Models
- Anthropic now bans users for sustained abuse toward Claude under new usage policy — chrisfirst · 2026-10-09
- LightOnOCR-3 open-sourced: one model for OCR, layout, image captions and chart extraction — IgorCarron · 2026-10-09
- OpenAI plans invisible watermarks on ChatGPT and Codex texts in the EU for compliance — emmanuelvivier · 2026-10-09
- Anthropic launches Claude Haiku 5.5 with 1M-token context, targeting high-volume tasks — emmanuelvivier · 2026-10-09
- OpenAI rolls out GPT-6 Intelligent UI: chatbot answers become interactive interfaces — emmanuelvivier · 2026-10-09
- Try LightOnOCR-3 on your hardest document: early user demos circulate — IgorCarron · 2026-10-09