Jev benchmarked: the only sub-second model that admits uncertainty across 16 tested

AlexKim · x · 2026-09-19

TypeSafe's Jev returns typed judgments with calibrated probabilities instead of prose: Choice picks an option, Score rates against ordered levels, Noul outputs the probability a yes/no condition holds. Independent questions share one state and run in parallel, so a twelve-check gate is a single call billed once.

The author measured 16 models on 150 passages (from 35 blogs, identifying the rewritten draft) across 2,400 calls, scoring expected calibration error:

The author also released a Claude Code hook to run the check.

Related event: 16-model calibration test: Jev fastest honest model but ranks 10th in accuracy(8 posts)→

Original post →

More from coding & agent

coding & agent channel →