JevBench open-sources a benchmark for typed decision models scoring intelligence, calibration and speed
sull · x · 2026-09-21
Developer fstandhartinger released JevBench on GitHub, a benchmark targeting "Jev-class" typed decision models.
- Method: the model receives a piece of state plus a bounded rubric and must return a typed answer with a probability for every option — no prose, no parsing.
- Systems measured: includes TypeSafe AI's Jev model, though the project states it is unaffiliated with the company.
- Current version: v1.2.3, with public results files, datasets (including notes on how the "hard tier" was built), and an interactive leaderboard at benchmarkheaven.com/jev-models.
- Scoring: the JevBench Score combines intelligence, calibration, speed and reliability.
Related event: JevBench v1.2 Released as Open-Source Decision Model Benchmark(2 posts)→
More from Research
- Brood War Bench: Codex Astra goes 18-0 while no model plays beyond beginner level — steipete · 2026-09-21
- JevBench v1.2: open-source LLM leaderboard weighting intelligence, calibration, speed, cost — airesearch12 · 2026-09-21
- MICCAI 2026 Tutorial to Explain Zeta Scaling Law Behind Medical AI Competition Rankings — PTenigma · 2026-09-21
- A 4B verifier locates hidden failures and boosts long-horizon agent reliability without retraining — teortaxesTex · 2026-09-21
- Stanford professor plans AI-assisted peer reviews of 5-6 papers monthly — anshulkundaje · 2026-09-21
- Researchers ask: why is there no AI-first scientific journal with AI-run peer review? — anshulkundaje · 2026-09-21