Hand-built anti-memorization benchmark: tiny Jev classifier matches Sonnet at ~150x lower cost
vesko_st · x · 2026-09-22
Developer veskost benchmarked TypeSafe AI's tiny Jev classifier and found it roughly on par with Claude Sonnet on classification tasks at orders of magnitude lower cost.
- On three standard reading-comprehension/commonsense benchmarks, Jev performed alongside Opus, topping CommonsenseQA and MMLU-CF and tying Opus on RACE-H, while being 150x cheaper and 10x faster on long-context questions.
- Suspecting training-data contamination, he hand-authored a fresh benchmark from a random Wikipedia article with distractors drawn from the text — Jev still landed alongside Sonnet and Opus.
- On older customer-service and negotiation datasets, Jev edged out Claude Haiku and sat slightly below Sonnet 5, though he cautions these aren't real B2B sales tasks.
He concludes small cheap classifiers like Jev have real uses and plans to adopt one in his product lightfld.
Related event: Tiny classifier Jev matches Claude Sonnet at ~150x lower cost(2 posts)→
More from Models
- Alignment backfires: model strips all faces from a deepfake detection dataset mid-task — generativist · 2026-09-22
- Grok 4.7 falls to #24 on Vals Index, down 5 points from Grok 4.6 — scaling01 · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22
- Jev reportedly does tensor logic under the hood: differentiable IF args, no wasted gen tokens — StewartalsopIII · 2026-09-22
- Multilingual Model Laya Trending on Hugging Face — convaiinnovations · 2026-09-22