Eve, a 0.6B open model, passes calibration: when it says 90%, believe it

colinmcnamara · x · 2026-09-24

One conclusion from the calibration thread: Eve, a 0.6B open model, narrowly passed — it reached 90% confidence on 115 of 1,599 questions, and only 12 of its 673 wrong answers claimed 90%, so its confidence statements are largely reliable. The thread also tested Jev, a hosted decision API: 96% right on the hard questions it kept, no GPU needed, but it is not deterministic and a third of its wrong answers on human labels still claimed 90%.

Related event: Calibration Tests: Small Models Often Claim 90% Confidence When Wrong(2 posts)→

Original post →

More from coding & agent

coding & agent channel →