.63 CLINC150 eval of Jev: excellent… · AGI Hunt

$0.63 CLINC150 eval of Jev: excellent calibration (ECE 0.025) but the fallback bucket can't be trusted

Artistic-Blood7501 · reddit · 2026-09-21

A self-funded eval of TypeSafe's Jev (jev-1.13.0) on full CLINC150 (4,500 in-scope + 1,000 out-of-scope utterances), 151-way choice, zero prompt tuning, zero request errors: 89.4% agreement, out-of-scope recall 83.8%/precision 89.2%, outstanding calibration (ECE 0.025, 96.2% of ≥0.9-confidence answers correct), 1.1s p50 latency, $0.63 total.

The key caveat: of 162 out-of-scope misses, 54 landed on specific intents at ≥0.9 confidence (TIME, DEFINITION, DATE pattern-matching), so neither the fallback option nor a confidence gate catches everything—treat the fallback bucket as weekly reading. Escalation table: gating at 0.7/0.8/0.9 catches 48%/61%/73% of errors at 93.7%/95.1%/96.2% accuracy. Results align with the pre-registered ASSAY-001 run. Full data in github.com/chr-kelly/jev-cookbook.

Original post →

More from Models

Models channel →