A model that refuses to talk: deep dive and hands-on test of TypeSafe's decision model Jev

colinmcnamara · x · 2026-09-21

Colin McNamara's long-form review of TypeSafe's Jev, a decision model that never writes prose. Key points: structured outputs fix the shape of chat output but not the mismatch — and a JSON confidence value is a generated estimate, not evidence of calibration. Jev opened early access Sept 15, 2026 alongside a $40M seed led by DCVC. His hands-on results: 94.0% sentiment, 88.2% news topics zero-shot; 143ms per question; ECE 0.09, underconfident on sentiment and overconfident at the top on news topics; an open 27B model on his own GPU calibrated at least as well. Actionable checklist: measure calibration on your own labeled data, set thresholds by mistake cost, pin the model version, and check accuracy and calibration as separate properties.

Related event: Practical playbook for shipping decision models(2 posts)→

Original post →

More from Models

Models channel →