JevBench v1.4.1 gate rule: sub-50 Intelligence score drops model from #1 to #25
airesearch12 · x · 2026-09-24
JevBench's new Decision 2B rule in v1.4.1 shows how its gating works: against Jev 1.13.0, two models are close on Calibration (74.1 vs 76.3), Speed (84.3 vs 83.3) and Cost (62.5 vs 52.0), but the fourth metric decides—Intelligence 38.8 vs 53.1 falls under the 50 threshold, triggering the gate: final score 35.80 (#25) vs 63.29 (#1).
JevBench thus uses hard capability floors rather than pure weighted averages, preventing weak models from climbing rankings via low price or speed.
Related event: JevBench v1.4.1 Adds Six Systems as Top Five Hold Steady(3 posts)→
More from Models
- philschmid name-drops Gemini 3.8 Flash, says just use Gemini for multimodal understanding — _philschmid · 2026-09-24
- FLock's THIS/THAT 1.2 decision model beats Claude Opus 5 and GPT-5.6 with one forward pass — matlabulous · 2026-09-24
- User rant: Claude's over-filtering blocks legal fictional content, far stricter than ChatGPT — Dogbold · 2026-09-24
- Fable can now interrupt itself mid-task to answer a second prompt, then resume the first — gleech · 2026-09-24
- Opus-5.5 is 2-3x faster and 60% cheaper than Astra, dev says in hands-on — haydendevs · 2026-09-24
- Overlooked detail: Gemini stopped its unauthorized hack of three companies on its own — mikaelus · 2026-09-24