Jev benchmarks: matches production classifiers on fixed-label tasks at ~100x lower cost
ivan_bezdomny · x · 2026-09-19
Developer drewdil benchmarked Typesafe AI's Jev against their production judge models and found it matches or beats them on fixed-label classification at a fraction of the cost.
- Sales-call transcript classification (32 call types): 35/36 correct on a synthetic set with 20 adversarial cases vs 34/36 for their current small-model classifier; zero label drift across three repeat runs; 100x cheaper per call.
- Knowledge-graph relation typing (fixed vocabulary): passed all four hard quality gates on a 47-case gold corpus, one case behind production, at 65x lower cost and 15x lower latency.
- Slack channel categorization (four buckets): 16/16 on the held-out set.
The takeaway: when the label list is defined by you and the evidence is in the input, Jev is a cheap, fast, consistent classifier; it's not suited for reasoning-heavy or recursive generation work.
More from Models
- Meta Muse + Jev screens 10,000 candidates to surface top 100 unanswered immunology questions — DeryaTR_ · 2026-09-19
- Braintrust adds Jev as a judge scorer: typed decisions at up to 193.6× speed and 444.6× lower cost — multiply_matrix · 2026-09-19
- GPU price hike hits even the 1080ti, as local LLM token-speed numbers circulate — HankYeomans · 2026-09-19
- I spent $3.40 on Jev in 24 hours: it will be Jev + LLMs, not Jev vs LLMs — gaganghotra_ · 2026-09-19
- GPT-6 Astra builds a flamethrower demo in Three.js from a screen recording — techartist_ · 2026-09-19
- Databricks: shifting just 20% of coding traffic to OSS models dramatically cuts inference spend — Yuchenj_UW · 2026-09-19