Six Real Workflows for the Jev Judgment Model: 24/24 Fact-Checks at 0.41s Median Latency
alexisgallagher · x · 2026-09-19
Isaac Flath plugged TypeSafe's new judgment model Jev into six real workflows, with a blog post and video including latency and accuracy evals:
- Fact-checking scripts: 24/24 correct like Gemini 3.5 Flash, but 0.41s median response vs 1.68s; also caught a 'half-hour call' that was actually call prep
- Feed ranking: replaced Gemini Flash as the LLM-as-a-judge for his multi-source news aggregator
- Locating the right text in PDFs
- Citation checking
- Clustering review notes
- Diagnosing agent failures
The pitch: swap slow, expensive LLM judges for a sub-second specialized judgment model — each use case comes with data against the previous setup, easy to replicate.
More from coding & agent
- Muse Opens Connector Platform: Developers Plug In APIs to Reach Agent Users — jffwng · 2026-09-19
- Jasper's open RL guide: shaping rewards to train a search agent end-to-end — simonguozirui · 2026-09-19
- Make Your Evals Better: More Edge Cases and Real User Inputs, Not Synthetic Slop — cephaloform · 2026-09-19
- Agent runs full SEO audit of Opendoor in 4.5 minutes, surfacing 500+ link opportunities — morganb · 2026-09-19
- AWS's Marc Brooker on the future of code review: LLMs plus automated reasoning — alexisgallagher · 2026-09-19
- Thorsten Ball demos agent picking its next command from shell history — IanArawjo · 2026-09-19