JevBench v1.2: open-source LLM leaderboard weighting intelligence, calibration, speed, cost
airesearch12 · x · 2026-09-21
Benchmark Heaven launched JevBench v1.2, an open-source leaderboard for "Jev-class" decision models. The main score is a geometric mean of intelligence, calibration, speed and cost, each weighted 25%. It offers alternative rankings emphasizing any single dimension, lets users re-configure weights for custom rankings, and supports filtering by hosting region, data confidentiality policies, open weights and price bases — released in response to criticism that prior comparison methodology was unclear.
Related event: JevBench v1.2 Released as Open-Source Decision Model Benchmark(2 posts)→
More from Models
- Burkov calls stealthy LLM-rival project Jev 'BS' over speed and calibration claims — burkov · 2026-09-21
- Moonshot and Tencent Hunyuan both building Flash models to target agent inference costs — TheZachMueller · 2026-09-21
- Dev Launches Made With Jev, a Free Directory Cataloging Demos, Tools and Skills for the New Model — Sea_Supermarket_5891 · 2026-09-21
- Which Model Actually Understands Reverse Engineering? Dev Seeks MCP Workflow for Ghidra and IDA Pro — obese_coder · 2026-09-21
- JEV's 'No Hallucination' Claim Under Fire: Same Prompts Give Different Probabilities Across Runs — FrankFelixAI · 2026-09-21
- Jev direct access now sits behind a waitlist after last week's surge — Creative-Drawer2565 · 2026-09-21