Jev vs generic model at same cost: 82.9% vs 58.8% on MMLU-Pro
simonguozirui · x · 2026-09-18
Developer ekzhang1 built a Jev-compatible public API running a comparable open model (Qwen3.6-35B-A3B, no reasoning), using only SGLang radix cache to preserve prefill reuse — generating 64 tasks in under a second.
Benchmarks: on MMLU-Pro, Jev scores 82.9% vs OpenJev's 58.8%; BoolQ calibration shows Jev slightly overconfident while OpenJev is underconfident. Conclusion: calibration is fine on both, but Jev is much smarter.
More from Models
- Jev playground shows inference and roundtrip latency; EU users pay 120ms extra — DanielLockyer · 2026-09-18
- Tesla publishes 12-month FSD Supervised safety report card — yunta_tsai · 2026-09-18
- ChatGPT co-inventor launches Jev model claiming 20-200x speed and 40-400x cost cuts — multiply_matrix · 2026-09-18
- GPT-6 Astra Cracks a Previously Unsolved 1918 German WWI Radio Message — iamfakhrealam · 2026-09-18
- Matt Holden: structured output is LLMs' real magic, now at 300ms and nearly free — holdenmatt · 2026-09-18
- Wrong Prediction, Right Answer: EMNLP paper shows a 'readout bottleneck' hides LLM reasoning — 机器之心 · 2026-09-18