What thousands of hours of agent runs reveal about LLM-as-judge: good at progress, bad at safety
hrishioa · x · 2026-09-21
hrishioa analyzed a few thousand hours of agentic runs using LLM-as-judge (Jev) and found: it's the best option he's tested for measuring progress and estimating completion; it's dangerous when used to detect harmful commands; and it's not good at catching models being lazy. The article includes actual prompts and results with detailed visualizations.
More from coding & agent
- Dev combines Jev and BaoCut into a local tool that finds video clips from one sentence — dotey · 2026-09-21
- claude-ops turns Claude Code into a business OS with 57 skills, 21 agents — tom_doerr · 2026-09-21
- Inside OpenAI's agentic software factory: Codex takes over, IDEs and pull requests fade — AxSaucedo · 2026-09-21
- Developer Gives AI Agent a Phone, Turns It Into a Personal Concierge — ethanniser · 2026-09-21
- Benchmarks show coding agents edit code they shouldn't in 35-65% of cases; prompt framing is the lever — RunAI_Coder · 2026-09-21
- Using a second LLM as a watchdog to catch coding agents faking success — Ascend-910 · 2026-09-21