Kev-9b and Decider-2b are accurate but wrongly claim 90% confidence on hard questions
colinmcnamara · x · 2026-09-24
Results from the calibration thread: Kev-9b and Decider-2b were the most accurate on human labels and answer in 5 ms, but on hard questions they claim 90% confidence on many wrong answers — use the answer, review everything. In the same thread, Eve, a 0.6B open model, narrowly passed: only 12 of 673 wrong answers claimed 90%, so its confidence claims are largely trustworthy.
Related event: Calibration Tests: Small Models Often Claim 90% Confidence When Wrong(2 posts)→
More from coding & agent
- Setting Astra's Thinking to 'Extra High' Makes It Over-Engineer Unit Tests — astralmatrix · 2026-09-24
- Free O'Reilly Book Offers a Pragmatic Framework for Scaling AI in Engineering Teams — blaizedsouza · 2026-09-24
- Zilliz CTO: agents make the enterprise data layer impossible to ignore — No_Engineer_1224 · 2026-09-24
- Running Android emulator + Chrome with 60fps streaming in a $0.072/hr cloud VM for always-on agents — cem2ran · 2026-09-24
- 10 agent reruns reached the right neighborhood, none reproduced the key observation — rohanpaul_ai · 2026-09-24
- Using the Jev model for offensive security: OPSC ranking and sensitive file detection in Mythic — dyn___ · 2026-09-24