Dev runs 16-model eval: Jev and Haiku tie at 0.121/0.122 on calibration error
AlexKim · x · 2026-09-19
A developer shares a systematic eval of 16 models, focusing on TypeSafe's Jev vs Haiku:
- Two AI-written draft conclusions contradicted each other: one claimed "Jev's confidence is usable, Haiku's is not," another said "Haiku was more accurate."
- Computing expected calibration error properly gave Jev 0.121 vs Haiku 0.122 — a tie on the very metric the product is built around.
- Multi-task results split: Haiku wins business categories 83.2 to 79.9, Jev wins commit types 50.0 to 42.0, prose dead-ties at 66.0.
- Key lesson: one task is not a benchmark, and AI-drafted eval conclusions shouldn't be trusted without recomputation.
More from Models
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19
- Empirical analysis confirms Claude Opus 5 shows abnormally dark base-model completions — Kyrannio · 2026-09-19
- US products quietly build on Chinese open-weight models as one firm cuts spend by ~100x — generativist · 2026-09-19
- Eval model Jev goes free on Vercel AI Gateway until Sept 25 — cramforce · 2026-09-19
- OpenAI confirms desktop Voice mode bug when starting new sessions, fix underway — juberti · 2026-09-19