Hands-on: TypeSafe's tiny Jev classifier goes head-to-head with Claude Haiku/Sonnet/Opus on text understanding
vesko_st · x · 2026-09-22
Ves Stoyanov (Lightfield) tested TypeSafe's newly released Jev, a compact "System One Choice" classifier that takes a passage, a question, and options and returns one answer. Wary that public benchmarks may be contaminated and stale, he re-ran everything under a fixed protocol and hand-wrote a corpus as a contamination hedge. Jev was benchmarked against Claude Haiku 4.5, Sonnet 5, and Opus 5 on CommonsenseQA, MMLU-CF, RACE-H, Grace and more, priced per 1,000 long-context questions. His takeaway: the purpose-built small classifier now holds its own on text understanding, making dedicated classifier models compelling again for routing-style tasks.
Related event: Tiny classifier Jev matches Claude Sonnet at ~150x lower cost(2 posts)→
More from Models
- Alignment backfires: model strips all faces from a deepfake detection dataset mid-task — generativist · 2026-09-22
- Grok 4.7 falls to #24 on Vals Index, down 5 points from Grok 4.6 — scaling01 · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22
- Jev reportedly does tensor logic under the hood: differentiable IF args, no wasted gen tokens — StewartalsopIII · 2026-09-22
- Multilingual Model Laya Trending on Hugging Face — convaiinnovations · 2026-09-22