LLM judges need judge-accuracy audits too — just like MiFID II requires for trading surveillance
sh_reya · x · 2026-09-25
Citing chapter 5 of "Evals for AI Engineers" by shreya and Hamel Husain: step 1 is measuring judge accuracy to estimate true success rates with imperfect judges. She points to MiFID II RTS 6, Article 13(6), which requires investment firms to review automated surveillance systems at least yearly for false positives/negatives — arguing an LLM judge should get the same audit every time its underlying model changes.
More from coding & agent
- Fighting agent scope drift: team adds Jev monitor to Hermes-Agent to keep agents on task — alexcovo_eth · 2026-09-25
- Connectome: open-source agent host that folds AI memories in layers like humans do — repligate · 2026-09-25
- Claude Code builds a full 3D survival game in one day: 2km island, 44k code-generated trees — majidmanzarpour · 2026-09-25
- Three states hold ComfyUI together: how SUCCESS, FAILURE, PENDING keep graph execution correct — Mahmoud_Zalt · 2026-09-25
- Agentic Software Factory: AI coding speed just moves the bottleneck downstream — Pavan_Belagatti · 2026-09-25
- Dev Builds Browser-Only Generator for OpenAI Strict JSON Schemas — Top-Coder-5852 · 2026-09-25