JevBench v1.5.5: top 10 unchanged, Jev still #1 on capability but only 3rd composite
airesearch12 · x · 2026-10-02
The JevBench leaderboard updated to v1.5.5: the top 10 on Capability Score is unchanged with Jev still first, but open competitors — Winnow, Cygnet, Surogate Rune — are neck and neck, and Jev ranks only 3rd on Composite Score. New in this version: a radar chart comparing models on typical Jev use cases, and the accompanying site (Benchmark Heaven) offers extensive filters for region (China/EU/US hosting and company origin), data-confidentiality policies, price bases across I/O blends, and open-weights-only views to help users pick the right model for their use case.
Related event: JevBench switches to Capability Score and adds custom cost controls(3 posts)→
More from Models
- Gemini 4 Argon posts lowest hallucination rate (15%) on AA-Omniscience benchmark — import_jmr · 2026-10-02
- User says Opus 5.5 shows no token savings, hits 5-hour limit in 3-4 messages — MarsupialFirst8617 · 2026-10-02
- llama.cpp adds Decision Models, expanding local inference capabilities — paf1138 · 2026-10-02
- Early Argon impressions: dev vibes-codes with it, says it shows no signs of benchmaxxing — cgarciae88 · 2026-10-02
- NVIDIA's Kumo-Tabular Tabular Foundation Model Trends on Hugging Face — nvidia · 2026-10-02
- Claude Sonnet 5.5, Grok 4.7 and GPT-6.1 Sol go live on Runware's OpenAI-compatible endpoint — aziz4ai · 2026-10-02