JevBench ranks models on typed decisions, weighing accuracy, calibration, latency and cost
rohanpaul_ai · x · 2026-09-20
A new benchmark, JevBench, targets models whose output is a bounded software decision rather than open-ended prose, following TypeSafe's Sept 15 release of Jev, which takes application state plus fixed choices and returns typed answers with probabilities.
- The composite score deliberately combines Intelligence, Calibration, Speed and Cost, since deployment can fail even with high raw accuracy
- Geometric mean prevents excellence on one axis from compensating for a weak one
- GPT-5.6 Luna beats Jev 1.13.0 on hard-case accuracy, yet Jev leads the composite thanks to latency, calibration and cost
- Authors caution this is evidence only for a narrow typed-decision workload, not general capability
More from Research
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20
- FutureHouse unveils "Bio Millennium Problems": hard-but-verifiable biology benchmarks for AI — _sholtodouglas · 2026-09-20
- Gated Recurrent Transformers: 3-layer recurrent model matches 12-layer GPT-2 with 63% fewer params — burny_tech · 2026-09-20
- Stanford builds virtual biotech run by 37,000 AI scientist agents, with real drug discovery results — Dr_Singularity · 2026-09-20
- A year after questioning LLM conjecture-solving, Cornell professor concedes AI surged ahead — burny_tech · 2026-09-20
- VLA-Replica: a $-low-cost real-world VLA benchmark where NVIDIA GR00T N1.7 matches π₀ with 50 demos — YuXiang_IRVL · 2026-09-20