Author to update evalstats library for complex LLM judge IRR reporting

IanArawjo · x · 2026-08-22

The author found that calculating IRR for an LLM judge is very complex for omnibus tests, concluding that the best solution is to use the evalstats library. They plan to adjust the library's judgealignment function to print the exact numbers and details required.

Original post →

More from coding & agent

coding & agent channel →