JevBench v1.4.1 gate rule: sub-50 Intelligence score drops model from #1 to #25

airesearch12 · x · 2026-09-24

JevBench's new Decision 2B rule in v1.4.1 shows how its gating works: against Jev 1.13.0, two models are close on Calibration (74.1 vs 76.3), Speed (84.3 vs 83.3) and Cost (62.5 vs 52.0), but the fourth metric decides—Intelligence 38.8 vs 53.1 falls under the 50 threshold, triggering the gate: final score 35.80 (#25) vs 63.29 (#1).

JevBench thus uses hard capability floors rather than pure weighted averages, preventing weak models from climbing rankings via low price or speed.

Related event: JevBench v1.4.1 Adds Six Systems as Top Five Hold Steady(3 posts)→

Original post →

More from Models

Models channel →