JevBench creator explains why open and API models get separate leaderboards: fairness
airesearch12 · x · 2026-10-11
The creator of JevBench responds to a Microsoft video, explaining why the benchmark splits open-weight and API-provided decision models onto two leaderboards: fairness. JevBench defines Jev-class models by capability, speed and cost combined, ranking on the Pareto frontier. But speed and cost can't be scored fairly across groups — API pricing and latency are business decisions, and a 'free' tier funded by VC subsidies could unfairly top a mixed ranking. Within the open-model group, all models run on the same hardware, so speed/cost comparisons are valid.
Related event: JevBench Author Explains Separate Rankings After Microsoft Citation Dispute(2 posts)→
More from Models
- Fact check: Alibaba Cloud is only retiring old model versions, not rival models — pstAsiatech · 2026-10-11
- Dead salmon fMRI case invoked to mock overinterpretation of LLM internals — IgorCarron · 2026-10-11
- Overzealous guardrails push devs from OpenAI's cyber model to GLM 5.3 for defensive auth code — yacineMTB · 2026-10-11
- Claude 3 Opus and Opus 5.5 already going full ASCII, user observes — repligate · 2026-10-11
- Xiaomi ships open-weight MiMo-V2.6-Pro and Flash with competitive coding scores — burkov · 2026-10-11
- Design Blogger Calls Out GrokBot: Email, Slack and Amazon Shopping, But No Real Use Case — talkaboutdesign · 2026-10-11