Pre-registered test of Jev's no-hallucination claim: type safety holds, calibration splits
NervousAd4440 · reddit · 2026-09-21
ASSAY-001 froze its protocol before API access and blind-scored by a third party: across 8,576 responses on Banking77 + CLINC150, Jev emitted zero out-of-schema values, validating the type-safety claim. But calibration was dataset-dependent: CLINC150 passed (ECE 0.0204) while Banking77 failed (ECE 0.0936) with systematic overconfidence in lower bins—schema-valid answers can still be wrong without warning, so threshold gates like act >0.9 must be fit per dataset. Two independent evals also went against Jev: Claude Haiku 4.5 beat it on 2,000-email phishing classification, and Jev alone didn't beat bge-m3 on 9,831 reranking pairs (only fusion did). The post links an awesome-jev repo that includes the losses.
More from Models
- Jev's eval abstraction maps 1:1 to autorubric paper from 8 months ago, researcher finds — deliprao · 2026-09-21
- Mollick: the most annoying part of long agentic tasks is language drift, not hallucinations — emollick · 2026-09-21
- No, Laya isn't capped at 512 tokens — it's ModernBERT with 8192-token configs — antoine_chaffin · 2026-09-21
- The catapult analogy: why AI models are jagged and robots face a deployment gap — lateinteraction · 2026-09-21
- Five error modes of frontier models used raw at the API level — gerardsans · 2026-09-21
- Qwen Image 2.1 license draws flak: 'the worst license yet' — switch2stock · 2026-09-21