JevBench creator explains why open and API models get separate leaderboards: fairness

airesearch12 · x · 2026-10-11

The creator of JevBench responds to a Microsoft video, explaining why the benchmark splits open-weight and API-provided decision models onto two leaderboards: fairness. JevBench defines Jev-class models by capability, speed and cost combined, ranking on the Pareto frontier. But speed and cost can't be scored fairly across groups — API pricing and latency are business decisions, and a 'free' tier funded by VC subsidies could unfairly top a mixed ranking. Within the open-model group, all models run on the same hardware, so speed/cost comparisons are valid.

Related event: JevBench Author Explains Separate Rankings After Microsoft Citation Dispute(2 posts)→

Original post →

More from Models

Models channel →