220K agent tool calls analyzed: model-watching-model Jev great at progress tracking, weak at security

hrishioa · x · 2026-09-21

Hrishi Olickel of Southbridge.AI evaluated typesafe.ai's Jev, a model that supervises other models, against 220K real tool calls (128K shell calls) from thousands of hours of agentic runs.

Key findings

Full prompts, data, and methodology (unified event logs + sentinel experiments on replays) are in the article.

Original post →

More from coding & agent

coding & agent channel →