JevBench v1.6.0 re-tests 92 models from scratch, new No. 1 emerges
airesearch12 · x · 2026-10-05
The author re-measured all 92 Jev-class models from scratch on brand-new sealed tests and released JevBench v1.6.0. Quyet-1.0-Large tops the Capability board at 81.7, ahead of deck-31B (77.6) and Jev 1.13.0 (76.5), reshuffling the top of the leaderboard.
More from Models
- Why don't modern LLMs know time has passed between messages? — dumierhan · 2026-10-06
- Reflection AI's new text model reportedly pretrained on ~24T tokens, multimodal version expected — nagpalchirag · 2026-10-06
- Early User Reports Anthropic's Opus 5.5 Fills Its Context Window Quickly — rickasaurus · 2026-10-06
- Viral Claude vs GPT Charts Mislead: Claude's "5x Value" Is Mostly Just Higher API Pricing — jdjohnson · 2026-10-06
- Rumor: Zhipu's next open source release GLM 5.5 may beat Claude Opus — bindureddy · 2026-10-06
- Liquid AI's d1 vision decision model matches GPT-6.1 Sol on 4 of 6 tasks at 19x-200x lower cost — JosephJacks_ · 2026-10-06