Full Jev eval released: 120 routes and 4,400 decision cases vs Qwen3 and Laya

paraschopra · x · 2026-09-19

paraschopra released full details of his Jev evaluation: a matched comparison from Sept 19, 2026 covering 120 navigation routes and 4,400 classification/decision cases against Laya and a local Qwen3-4B decision model. Findings: Jev leads on most tasks; Laya's English is not an overall improvement over the local prototype; none reliably solves their WikiRouter setup. Combined with earlier MMLU comparisons, Jev looks 30B-scale with 300ms latency and a telling 0% on relational choice benchmarks.

Related event: Mystery probability-only model Jev sparks flurry of community benchmarks(7 posts)→

Original post →

More from Models

Models channel →