JevBench v1 puts nine typed-decision models head-to-head across 242 decisions
airesearch12 · x · 2026-09-19
- The author released JevBench v1, an open benchmark for 'Jev-class' typed decision models: given a piece of state and a bounded rubric, models return a typed answer with a probability — no prose, no parsing.
- The suite runs 242 decisions across six families, scored on five axes: smart (accuracy), cheap (cost per 1,000 decisions), fast (end-to-end latency incl. network), reliable (probability calibration, rephrasing stability, schema adherence), and open (weights and licence).
- Results: official Jev 1.13.0 scores 96.3% at $0.027/1k and 0.65s latency; GPT-5.6 Luna tops accuracy at 97.1% but costs $0.176/1k; the best open rebuild, openjev-sglang, hits 95.5%.
- Part of the Benchmark Heaven project, with an MIT-licensed harness, public split, code and results on GitHub; explicitly not affiliated with TypeSafe AI.
More from Models
- Codex usage limits finally reset after days of users hitting exhausted quotas — CtrlAltDwayne · 2026-09-19
- Rumored Opus 5.2 generates a full 5-minute Titanic film in pure JavaScript with three.js — dotey · 2026-09-19
- Model Remembers 'cranberry-42' Across App Restarts in Persistent-Session Test — RileyRalmuto · 2026-09-19
- Free models aren't the real problem: studies show hallucinations persist in SOTA LLMs — AryHHAry · 2026-09-19
- Why do AI models lack personality? Safety training and attachment fears, explained — alexisgallagher · 2026-09-19
- New model JEV is free on OpenRouter, Vercel, Cloudflare and others, at ~$0.042/M input tokens — airesearch12 · 2026-09-19