Jev's calibration error measured at ~0.09: right just as often, still off about how sure
colinmcnamara · x · 2026-09-21
Against human labels across two prompts, TypeSafe's decision model Jev showed an expected calibration error of about 0.09 — underconfident on sentiment, overconfident at the top on news topics. It is right just as often, but still off about how sure it is.
More from Models
- YOCO back in spotlight: blog breaks down cross-layer KV sharing in DeepSeek-V4.1-Flash and Gemma 4 — donglixp · 2026-09-21
- Gemini 4.0 Rumored Next: Yearly Pro Updates, Monthly Flash Releases Expected — haider1 · 2026-09-21
- SemiAnalysis' Dylan Patel: Chinese labs going closed-source, 'open is dying quickly' — ns123abc · 2026-09-21
- DeepSeek Web Output Style Reportedly Lobotomized by Safe Harbor RLHF — Old_Let6328 · 2026-09-21
- JEV opens to all with $5 free credits; $0.042 input and free output pricing — op7418 · 2026-09-21
- ChatGPT 20x Max users hit tighter limits, suspect compute favors government and API customers — BopSupreme · 2026-09-21