Jev vs generic model at same cost: 82.9% vs 58.8% on MMLU-Pro

simonguozirui · x · 2026-09-18

Developer ekzhang1 built a Jev-compatible public API running a comparable open model (Qwen3.6-35B-A3B, no reasoning), using only SGLang radix cache to preserve prefill reuse — generating 64 tasks in under a second.

Benchmarks: on MMLU-Pro, Jev scores 82.9% vs OpenJev's 58.8%; BoolQ calibration shows Jev slightly overconfident while OpenJev is underconfident. Conclusion: calibration is fine on both, but Jev is much smarter.

Original post →

More from Models

Models channel →