Flowchart says LLM judges need at least 15 labels and IRR around 0.40

IanArawjo · x · 2026-07-29

A decision flowchart proposes when LLM judges are worth using in mixed human-AI evaluation studies, and when researchers should fall back to classical human-only statistics.

Key points:

Related event: LLM Judge Flowchart Sparks Debate on Power Analysis(2 posts)→

Original post →

More from Research

Research channel →