LLM as a judge: How do you trust the judge?
maylad31 · reddit · 2026-08-26
The author highlights the instability of using LLMs as judges, noting that changing field order can yield different results. Recommendations include evaluating consistency, preferring well-defined categories or clear rubrics, and avoiding arbitrary numerical scores with simple prompts. The post emphasizes not blindly trusting LLM judges and discusses methodologies for evaluating the evaluator itself.
More from coding & agent
- Computer use works far better behind async channels like iMessage, says illscience — illscience · 2026-08-26
- Scored my agent workflow on portability: 6/8, and the four layers that decide migration pain — SaschaFromWhaaat_ai · 2026-08-26
- Dev wired Claude to autonomously ship hourly changes to a 17-year-old SaaS in production — mhmazur · 2026-08-26
- Claude's web search stops at 20-30 results — user asks how to force exhaustive scraping — AskLendnow · 2026-08-26
- An agent that remembers everything has bad memory — how to maintain useful memory — victorialslocum · 2026-08-26
- Agent Sandbox Usage: Receipt URL Becomes Critical Verification Gate — Common_Dream9420 · 2026-08-26