ZenML docs: judge agent sessions with TypeSafe's jev model, ~270ms per call
strickvl · x · 2026-09-24
ZenML published full docs for Kitaru judge evaluations: TypeSafe's jev takes JSON state and typed questions, returning yes/no probabilities, a chosen label with confidence, or a position on ordered levels, one result per question.
The docs compare three approaches — deterministic evaluators (free, code rules only), the typed model judge (270ms per call in exploratory runs, can catch invented facts), and hand-written LLM judges (flexible but cost/repeatability depend on model and prompt) — and when to use each.
More from coding & agent
- "This Is Too Slow, Make It Faster" — Dev's One-Liner Actually Optimized His AI Code — DanielLockyer · 2026-09-24
- SwiftFairy launches: on-device Mac app that reviews your coding agent's Swift code — JordanMorgan10 · 2026-09-24
- Stripe's Link Wallet Tops 300M Consumers, Partners With Muse for Agentic Buying — jeff_weinstein · 2026-09-24
- Zuckerberg Put Cameras in His MMA Gym So an Agent Could Coach Him Between 1-Minute Rounds — victor_explore · 2026-09-24
- LangChain's Interrupt keynote ships LangSmith Engine v2, Deep Agents 0.8, fine-tuning and more — LangChain · 2026-09-24
- Runable launches Scheduled Tasks: agents that do the work, not just remind you — SimplyAnnisa · 2026-09-24