16-model calibration test: open-weight models almost never admit uncertainty

AlexKim · x · 2026-09-19

A calibration test across 16 models measured how often each puts answers in the ambiguous 0.35–0.65 middle: Sonnet 5 leads at 41.3%, Fable 5.1 at 36.7%, Jev 34.7%, Opus 5 at 23.3%, then a cliff — nothing else clears 7.3%. GPT-5.5 is the best of its family at 11.3%, still below the group, and no open-weight model clears it, with kimi-k2.5 and deepseek-v4-pro at just 0.7%.

Related event: 16-model calibration test: Jev fastest honest model but ranks 10th in accuracy(8 posts)→

Original post →

More from Models

Models channel →