JevBench v1.4 adds 308 evolving sealed tasks to stop benchmaxxing, now covers 70+ models
airesearch12 · x · 2026-09-23
JevBench has released v1.4, now covering 70+ models with the top 5 being Jev, JevK5, Hopper, Winnow-12B Q8, and reflex 4B. The update focuses on anti-benchmaxxing:
- Evolving sealed test set: 308 new hard tasks that keep evolving, supplying 20% of the Intelligence score, to prevent public-set saturation and training-data leakage.
- Calibration: sealed results are blended into calibration at a moderated weight; a k=1 penalty targets public-to-sealed gaps above 25 pp, rewarding genuine generalization.
- Equal-weight harmonic mean across four axes.
- Speed and cost gating: brilliant models still need Jev-class speed and cost, or they get gated down.
- API transparency: the leaderboard now flags whose API endpoints received held-out test item texts; answers are never shared.
Related event: Satirical JevBench v1.4 now ranks over 70 models, with 'Jev' still on top(3 posts)→
More from Models
- 10K-run test: LLM hits 99.91% on 30-digit multiplication, 99.8% on division — srchvrs · 2026-09-23
- Heavy research user: Opus 5.5 is the first Claude I want as my daily driver since 4.7 — jxnlco · 2026-09-23
- Ben Todd: Google Astra Suddenly Gained Car-Driving Ability Without Specialist Training — ben_j_todd · 2026-09-23
- NVIDIA releases Nemotron 3 Diarization model handling up to 8 overlapping speakers with 100M params — NVIDIAAI · 2026-09-23
- Dev slams Anthropic's Opus 5.5 safety checks for flagging basic code reviews — evilsocket · 2026-09-23
- Deep conversations with frontier models turn into incomprehensible AI-to-AI jargon, observer warns — erikphoel · 2026-09-23