Pre-registered test of Jev's no-hallucination claim: type safety holds, calibration splits

NervousAd4440 · reddit · 2026-09-21

ASSAY-001 froze its protocol before API access and blind-scored by a third party: across 8,576 responses on Banking77 + CLINC150, Jev emitted zero out-of-schema values, validating the type-safety claim. But calibration was dataset-dependent: CLINC150 passed (ECE 0.0204) while Banking77 failed (ECE 0.0936) with systematic overconfidence in lower bins—schema-valid answers can still be wrong without warning, so threshold gates like act >0.9 must be fit per dataset. Two independent evals also went against Jev: Claude Haiku 4.5 beat it on 2,000-email phishing classification, and Jev alone didn't beat bge-m3 on 9,831 reranking pairs (only fusion did). The post links an awesome-jev repo that includes the losses.

Original post →

More from Models

Models channel →