What thousands of hours of agent runs reveal about LLM-as-judge: good at progress, bad at safety

hrishioa · x · 2026-09-21

hrishioa analyzed a few thousand hours of agentic runs using LLM-as-judge (Jev) and found: it's the best option he's tested for measuring progress and estimating completion; it's dangerous when used to detect harmful commands; and it's not good at catching models being lazy. The article includes actual prompts and results with detailed visualizations.

Related event: Analysis of 220K Agent Tool Calls: LLM Judges Good at Tracking Progress, Weak at Catching Attacks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →