A model that refuses to talk: deep dive and hands-on test of TypeSafe's decision model Jev
colinmcnamara · x · 2026-09-21
Colin McNamara's long-form review of TypeSafe's Jev, a decision model that never writes prose. Key points: structured outputs fix the shape of chat output but not the mismatch — and a JSON confidence value is a generated estimate, not evidence of calibration. Jev opened early access Sept 15, 2026 alongside a $40M seed led by DCVC. His hands-on results: 94.0% sentiment, 88.2% news topics zero-shot; 143ms per question; ECE 0.09, underconfident on sentiment and overconfident at the top on news topics; an open 27B model on his own GPU calibrated at least as well. Actionable checklist: measure calibration on your own labeled data, set thresholds by mistake cost, pin the model version, and check accuracy and calibration as separate properties.
Related event: Practical playbook for shipping decision models(2 posts)→
More from Models
- MiMo near-SOTA on DeepSWE with just ~$2.6M RL run: will data cost more than training? — my_cat_can_code · 2026-09-21
- humansand's Persimmon model learns to share info gradually like humans, with Trickle Test — niloofar_mire · 2026-09-21
- Why yes/no answers are fast for LLMs: output tokens dominate latency — tinyfool · 2026-09-21
- ChatGPT reportedly removes free-tier chat limits, offering unlimited text chats — nikola_mr64990 · 2026-09-21
- Users say top-tier Astra is too costly, hope GPT-6 fixes token economics — CtrlAltDwayne · 2026-09-21
- TypeSafe's JEV model fully open with $5 free credit, powers 1-second 3D scene generation — tinyfool · 2026-09-21