Analysis of 220K Agent Tool Calls: LLM Judges Good at Tracking Progress, Weak at Catching Attacks

Hrishi Olickel of Southbridge.AI analyzed thousands of hours of agentic runs and about 220,000 tool calls to evaluate the LLM-based supervisor Jev, finding it the best method he has tested for measuring progress and completion, but unreliable for catching shirking and defending against attacks.

2026-09-21 ~ 2026-09-21 · 2 related posts