Calibration Tests: Small Models Often Claim 90% Confidence When Wrong

Calibration benchmarks show Kev-9b and Decider-2b give wrong answers with 90% stated confidence on hard questions, while the 0.6B open model Eve passes: when it claims 90% confidence, it is usually right.

2026-09-24 ~ 2026-09-24 · 2 related posts