16-model calibration test: open-weight models almost never admit uncertainty
AlexKim · x · 2026-09-19
A calibration test across 16 models measured how often each puts answers in the ambiguous 0.35–0.65 middle: Sonnet 5 leads at 41.3%, Fable 5.1 at 36.7%, Jev 34.7%, Opus 5 at 23.3%, then a cliff — nothing else clears 7.3%. GPT-5.5 is the best of its family at 11.3%, still below the group, and no open-weight model clears it, with kimi-k2.5 and deepseek-v4-pro at just 0.7%.
More from Models
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19
- Empirical analysis confirms Claude Opus 5 shows abnormally dark base-model completions — Kyrannio · 2026-09-19
- US products quietly build on Chinese open-weight models as one firm cuts spend by ~100x — generativist · 2026-09-19
- Eval model Jev goes free on Vercel AI Gateway until Sept 25 — cramforce · 2026-09-19