Dev runs 16-model eval: Jev and Haiku tie at 0.121/0.122 on calibration error

AlexKim · x · 2026-09-19

A developer shares a systematic eval of 16 models, focusing on TypeSafe's Jev vs Haiku:

Related event: 16-model calibration test: Jev fastest honest model but ranks 10th in accuracy(8 posts)→

Original post →

More from Models

Models channel →