Model 'Jev' shows well-calibrated probabilities: 1.74pp average calibration error across benchmarks
hackgoofer · x · 2026-09-17
N8Programs revealed that model Jev outputs well-calibrated probabilities: an expected calibration error of just 1.74 percentage points averaged across benchmarks, ranging from 0.26pp at best to 6.96pp at worst.
The project team said community response has been strong and opened access—users can message the #meme channel in their Discord and DM a screenshot with an email to get in.
More from Models
- Bio speaker still mocks ChatGPT hallucinations; author asks if they even used deep research — zebird0 · 2026-09-17
- Flagship models cost 2-4x more for marginal gains: GPT-6 Astra at $3.94/task vs $1.03 — arena · 2026-09-17
- Ask GPT-5.2 and Claude Opus 4.6 to 'Be the Null' and They Output Zero Bytes, 30/30 — rayanpal_ · 2026-09-17
- Codex hits 20M users as free resets reportedly end ahead of DevDay — brandon_galang · 2026-09-17
- GLM 5.3 lands in Brave Nightly, rivaling GPT 5.6 Sol and Grok 4.6 — gnukeith · 2026-09-17
- Benchmark shows GPT 5.6 Luna matches GPT 5.5 intelligence at far better value — nickbaumann_ · 2026-09-17