JevBench v1.4.1: Jev-Omni edges hard tier 75.0% vs 74.1%, but trails badly on calibration
airesearch12 · x · 2026-09-24
Third-party benchmark site Benchmark Heaven published frozen JevBench v1.4.1 results:
- New Jev-Omni (v1.4.1) narrowly beats Jev 1.13.0 on the hard tier: 75.0% vs 74.1%.
- It loses everywhere else: sealed set 32.1% vs 36.7%, calibration 64.1 vs 76.3, composite score 51.34 (#7) vs 63.29 (#1).
- Takeaway: the models are close on answers but far apart on confidence — answer accuracy doesn't equal reliability.
Related event: JevBench v1.4.1 Adds Six Systems as Top Five Hold Steady(3 posts)→
More from Models
- JevBench splits capability from cost and speed in new plots, refusing to blur them into one score — airesearch12 · 2026-09-24
- JevBench v1.4.1 gate rule: sub-50 Intelligence score drops model from #1 to #25 — airesearch12 · 2026-09-24
- Codex user asks whether usage resets carry over after upgrading to 20x plan — justalexoki · 2026-09-24
- ChatGPT reportedly gives free users unlimited GPT-5.6 Luna text chats — hey_abusiddik · 2026-09-24
- JevBench: DeepSeek V4.1 Flash outscores leader at 1/15th the cost per decision — airesearch12 · 2026-09-24
- Altman claims OpenAI model solved Navier-Stokes, a Millennium Prize problem — victor_explore · 2026-09-24