Getting an LLM judge from garbage to 27/30 human agreement: an evals workshop writeup

tejaskumarlol · reddit · 2026-10-08

A hands-on recap of an AI Engineer World's Fair workshop on building a reliable LLM-as-judge, taking it from garbage to 27/30 agreement with human verdicts.

Author works at IBM; the demo's retrieval used their open-source OpenRAG.

Original post →

More from coding & agent

coding & agent channel →