LLM as a judge: How do you trust the judge?

maylad31 · reddit · 2026-08-26

The author highlights the instability of using LLMs as judges, noting that changing field order can yield different results. Recommendations include evaluating consistency, preferring well-defined categories or clear rubrics, and avoiding arbitrary numerical scores with simple prompts. The post emphasizes not blindly trusting LLM judges and discusses methodologies for evaluating the evaluator itself.

Original post →

More from coding & agent

coding & agent channel →