Five exercises find four defects in the author's own LLM-judge eval system
alexpran · reddit · 2026-09-15
Applying Dan Luu's 'what's wrong with this benchmark?' method to his own LLM-judge evals, the author surfaces four defects: non-unanimous rates swing from 2/21 to 6/21 across 16 identical replicate runs; a 17→144 case suite breaks the unchanged fail-threshold rule; precision computed on half the pipeline for two weeks looked plausible the whole time; and two sensible prompt edits dropped recall from 0.80 to 0.50 or narrowed the rule they meant to widen. Answers and run files in the first comment.
More from coding & agent
- Aholo Lux3D turns text or one photo into a production-ready 3D model in 20 seconds — JaynitMakwana · 2026-09-15
- Developer Shares Hands-On Use of a Claude Code Persistent Memory Skill — doodlestein · 2026-09-15
- Hermes Newswire, a zero-token RSS news ticker plugin, merged into NousResearch's official catalog — Teknium · 2026-09-15
- vphone-cli lets AI agents control a fully virtualized iPhone on Apple Silicon, +633 stars in a day — JiliJeanlouis · 2026-09-15
- Codex users report built-in deep-research skill vanishing from CLI overnight — Hot_Independence5160 · 2026-09-15
- DeepSeek's open-source dsh harness makes every part of a coding agent a swappable plugin — alex_verem · 2026-09-15