Testing TypeSafe's Jev classifier: calibration check costs under $4, results mixed
RexDouglass · x · 2026-10-07
Dylan Black stress-tests Jev, TypeSafe's new "System One" classifier model (LLMs = slow general System Two; Jev = fast specialized System One).
- Jev bolts a classifier onto a pretrained transformer: give it context plus a multiple-choice spec like jev("2+2=?", "3 : int, 4 : int, 5 : int") and get a typed answer with a probability distribution
- Extremely cheap: the entire experiment series cost under $4.00
- The key question is whether those probabilities are calibrated — the author checks against well-understood questions and concludes Jev is poorly calibrated
More from Models
- User puzzled: ChatGPT keeps shilling an unknown third-party service — doodlestein · 2026-10-07
- Surrogate Rune v3 and Xor 26B-A4B join JevBench — airesearch12 · 2026-10-07
- Mistral Large 4 doubles its AI Index score to 38, open weights shipping this month — shensi · 2026-10-07
- Intel researcher: I'd never hire someone who thinks temp=0 guarantees LLM determinism — JFPuget · 2026-10-07
- New Mistral Is API-Ready but Not Open-Weight: Should It Lead Open-Weight Boards? — CharlotteHase · 2026-10-07
- AI usage inflects exponentially past capability thresholds — is research next? — menhguin · 2026-10-07