JevBench hits HN frontpage, critics allege Jev-class model is a thin Qwen wrapper
airesearch12 · x · 2026-09-23
JevBench, a reproducible benchmark for typed decision models, hit the HN frontpage (103 points, 25 comments) and drew sharp skepticism.
- Design: Jev-class models return bounded choices and probabilities instead of text; the bench weighs accuracy, latency and price together, scoring chance-corrected Intelligence, Calibration, Speed and Cost over 534 English decisions.
- Leaderboard: Jev 74.4, SemIf 73.1, djev 73.0, Winnow-12B Q8 71.2, reflex 4B 70.3.
- Openness: MIT-licensed harness, public items, frozen artifacts, public per-task outcomes, plus two no-signup demos.
- HN pushback: A top comment notes Jev has $40m funding after 2 years in stealth, yet performs on par with SemIf—built in days on raw Qwen, runs in-browser, and costs half as much—and consistently thinks it's Qwen, suggesting a possible scam.
- Disclosed limitations: English-only, single German server latency, ×2 demo latency adjustment, 1-point gaps may be noise.
More from Fun
- Jev clone wave: dozens of open-source alternatives appear within a week of launch — airesearch12 · 2026-09-23
- Unscripted Claude Opus demo: pixel-art solar system synced to Holst's The Planets — technollama · 2026-09-23
- A week of Jev, GPT sol 6, opus 5.5: netizens can't keep track of model releases anymore — malliktwts · 2026-09-23
- AI Twitter Debate: Stacking NVIDIA B300s Can't Beat Hackers Who Started Young, JEE Blamed for Killing Talent — itsOmSarraf_ · 2026-09-23
- Major model release gap shrinks from 73 days in 2023 to 18 days in 2026 — MickeySteamboat · 2026-09-23
- \paragraph{} Is the Vestigial Organ of All Semi-Scholarly AI-Generated Prose — aryaman2020 · 2026-09-23