JevBench v1.6.0 Re-tests 92 Models on Sealed Sets, Leaderboard Reshuffled
airesearch12 · x · 2026-10-05
The author re-measured all 92 Jev-class models from scratch on brand-new sealed tests, and the top of the leaderboard looks completely different: Quyet-1.0-Large takes #1 on Capability (81.7), ahead of deck-31B (77.6) and Jev 1.13.0 (76.5).
v1.6.0 is the most contamination-resistant release yet: a fresh sealed set nobody has seen each release, hosted APIs get items never sent to any provider, and items retire after limited use — so scores reflect real ability, not memorised test items.
More from Models
- Subscription "API value" math is inflated by token pricing, researcher points out — JeremyNguyenPhD · 2026-10-06
- Why don't modern LLMs know time has passed between messages? — dumierhan · 2026-10-06
- Reflection AI's new text model reportedly pretrained on ~24T tokens, multimodal version expected — nagpalchirag · 2026-10-06
- Early User Reports Anthropic's Opus 5.5 Fills Its Context Window Quickly — rickasaurus · 2026-10-06
- Viral Claude vs GPT Charts Mislead: Claude's "5x Value" Is Mostly Just Higher API Pricing — jdjohnson · 2026-10-06
- Rumor: Zhipu's next open source release GLM 5.5 may beat Claude Opus — bindureddy · 2026-10-06