Jev's calibration error measured at ~0.09: right just as often, still off about how sure

colinmcnamara · x · 2026-09-21

Against human labels across two prompts, TypeSafe's decision model Jev showed an expected calibration error of about 0.09 — underconfident on sentiment, overconfident at the top on news topics. It is right just as often, but still off about how sure it is.

Original post →

More from Models

Models channel →