Full Jev eval released: 120 routes and 4,400 decision cases vs Qwen3 and Laya
paraschopra · x · 2026-09-19
paraschopra released full details of his Jev evaluation: a matched comparison from Sept 19, 2026 covering 120 navigation routes and 4,400 classification/decision cases against Laya and a local Qwen3-4B decision model. Findings: Jev leads on most tasks; Laya's English is not an overall improvement over the local prototype; none reliably solves their WikiRouter setup. Combined with earlier MMLU comparisons, Jev looks 30B-scale with 300ms latency and a telling 0% on relational choice benchmarks.
Related event: Mystery probability-only model Jev sparks flurry of community benchmarks(7 posts)→
More from Models
- Moonshot's Kimi K3 lands on Amazon Bedrock with 1M-token context and prompt caching — emmanuelvivier · 2026-09-19
- ChatGPT Searched 74 Websites to Answer a Daylight Savings Question — dioscuri · 2026-09-19
- DeepSeek V5 Leak: Imminent Launch Could Match Top Closed Models, Staying Open-Weight — inductionheads · 2026-09-19
- Power user burns 1bn tokens a day at 98.65% cache hits, weighing GLM-5.3 vs DeepSeek after Claude Code weekly limits — julianharris · 2026-09-19
- Open models now outspend OpenAI on Vercel gateway, taking 78% of token volume — soumitrashukla9 · 2026-09-19
- Minimax-H3 dataset trending on Hugging Face under MIT license — stablediffusiontutorials · 2026-09-19