Flowchart says LLM judges need at least 15 labels and IRR around 0.40
IanArawjo · x · 2026-07-29
A decision flowchart proposes when LLM judges are worth using in mixed human-AI evaluation studies, and when researchers should fall back to classical human-only statistics.
Key points:
- First ask whether AI judges are worth the trouble at all: large datasets and a separate tuning set make them more viable.
- Randomly sample items for human labels, keeping the judge aligned on a tuning set.
- The flowchart recommends a minimum labeled sample size of Nlabl ≥ 15 and an inter-rater reliability threshold of about IRR ≥ 0.40 before trusting judge + human data.
- If the data are too small, or if the metric is easy to code automatically, the advice is to avoid LLM judges and revert to classical statistics on human-only or coded metrics.
- Even with high IRR, the figure stresses that this is not a license to skip PPI correction; researchers should run the appropriate PPI-corrected test and report judge model, prompt/setup, alignment metric, and results.
Related event: LLM Judge Flowchart Sparks Debate on Power Analysis(2 posts)→
More from Research
- Engineer Debunks Kimi K3 Memory Claims: Small State ≠ Flash Offload — AccBalanced · 2026-07-30
- Discussion: Why hasn't anyone built a neural network to detect AI text? Image detection has research papers — emeka_boris · 2026-07-30
- Agents still struggle with mathematical work: Codex spirals into 'proof certificates' and inventories — doodlestein · 2026-07-30
- TorchSpec Enables Disaggregated Speculative Decoding Training at Scale — zhyncs42 · 2026-07-30
- Compute Surge: 10 Major Scientific Breakthroughs AI Could Unlock by 2028 — Annual_Judge_7272 · 2026-07-30
- Inside SOTA Deep Research: Native Model Training and 150 Sub-Agents — SimonShaoleiDu · 2026-07-30