Kev-9b and Decider-2b are accurate but wrongly claim 90% confidence on hard questions

colinmcnamara · x · 2026-09-24

Results from the calibration thread: Kev-9b and Decider-2b were the most accurate on human labels and answer in 5 ms, but on hard questions they claim 90% confidence on many wrong answers — use the answer, review everything. In the same thread, Eve, a 0.6B open model, narrowly passed: only 12 of 673 wrong answers claimed 90%, so its confidence claims are largely trustworthy.

Related event: Calibration Tests: Small Models Often Claim 90% Confidence When Wrong(2 posts)→

Original post →

More from coding & agent

coding & agent channel →