Jev Aces 40 Math Problems in 2 Seconds for $0.0008, Matching GPT-5.4 Accuracy
xeophon · x · 2026-09-23
Multiple benchmarks show the new Jev model punching far above its cost:
- CEMC math contests (Canada, grades 9-11, 40 multiple-choice): Jev scored 60%, finishing all 40 questions in 2 seconds at $0.0008; GPT-5.4 with reasoning off scored 59% at $0.017 (20x the cost), and 5.4 nano scored 34% at $0.0014 (1.6x).
- GPQA Diamond: Jev hit 77%, comparable to GPT-5.5 with reasoning off or Fable 5 at low reasoning — completing all 198 questions in 3 seconds for $0.007.
The standout is extreme speed and cost efficiency: near-parity accuracy at 1-2 orders of magnitude lower inference cost.
More from Models
- Beff Jezos jokes Opus 5.5 is 'post-slop', freeing readers from AI sludge prose — beffjezos · 2026-09-23
- Opus 5.5 One-Shots a Full Prince of Persia Level with NPCs, Sound and Music — iannuttall · 2026-09-23
- The Full Prompt Behind Opus 5.5's One-Shot Prince of Persia Level — iannuttall · 2026-09-23
- Model release fatigue: devs say gains are marginal, but some argue new models clearly outpace old ones — Rasmic · 2026-09-23
- Follow-up: Opus 5.5 roughly on par with Astra and Fable 5.1, no clear winner — AaronBergman18 · 2026-09-23
- Opus 5.5 tentatively the world's smartest model, though slightly behind on world knowledge — AaronBergman18 · 2026-09-23