$0.63 CLINC150 eval of Jev: excellent calibration (ECE 0.025) but the fallback bucket can't be trusted
Artistic-Blood7501 · reddit · 2026-09-21
A self-funded eval of TypeSafe's Jev (jev-1.13.0) on full CLINC150 (4,500 in-scope + 1,000 out-of-scope utterances), 151-way choice, zero prompt tuning, zero request errors: 89.4% agreement, out-of-scope recall 83.8%/precision 89.2%, outstanding calibration (ECE 0.025, 96.2% of ≥0.9-confidence answers correct), 1.1s p50 latency, $0.63 total.
The key caveat: of 162 out-of-scope misses, 54 landed on specific intents at ≥0.9 confidence (TIME, DEFINITION, DATE pattern-matching), so neither the fallback option nor a confidence gate catches everything—treat the fallback bucket as weekly reading. Escalation table: gating at 0.7/0.8/0.9 catches 48%/61%/73% of errors at 93.7%/95.1%/96.2% accuracy. Results align with the pre-registered ASSAY-001 run. Full data in github.com/chr-kelly/jev-cookbook.
More from Models
- Dev buys a Meta coding subscription for its 'excellent model, crazy quota, low price' — intellectronica · 2026-09-21
- Rumored Opus 5.2/5.5 outputs circulate; execs reportedly expect taste gap to close — teortaxesTex · 2026-09-21
- Astra for prose, Fable for long-horizon work: a developer's multi-model division of labor, with Pi as best harness — seatedro · 2026-09-21
- Zero-shot embedding classifiers: prototyping superpower or lazy black box? — antoine_chaffin · 2026-09-21
- Bespoke-Nimble-9B, a Qwen3.5-9B LoRA for evidence-grounded text classification, trends on Hugging Face — bespokelabs · 2026-09-21
- Sophia Yang Gets Jev Access, Runs Small RL Experiment With Fireworks AI — sophiamyang · 2026-09-21