JevBench v1.4 follow-up link: details of the anti-benchmaxxing methodology
airesearch12 · x · 2026-09-23
A follow-up post linking to the JevBench v1.4 methodology details: 308 evolving sealed tasks (20% of the Intelligence score), a k=1 penalty for public-to-sealed gaps above 25 pp, equal-weight harmonic mean across four axes, speed/cost gating, and API endpoint transparency for held-out items.
Related event: Satirical JevBench v1.4 now ranks over 70 models, with 'Jev' still on top(3 posts)→
More from Models
- 10K-run test: LLM hits 99.91% on 30-digit multiplication, 99.8% on division — srchvrs · 2026-09-23
- Heavy research user: Opus 5.5 is the first Claude I want as my daily driver since 4.7 — jxnlco · 2026-09-23
- Ben Todd: Google Astra Suddenly Gained Car-Driving Ability Without Specialist Training — ben_j_todd · 2026-09-23
- NVIDIA releases Nemotron 3 Diarization model handling up to 8 overlapping speakers with 100M params — NVIDIAAI · 2026-09-23
- Dev slams Anthropic's Opus 5.5 safety checks for flagging basic code reviews — evilsocket · 2026-09-23
- Deep conversations with frontier models turn into incomprehensible AI-to-AI jargon, observer warns — erikphoel · 2026-09-23