Decision Index 0.1 benchmarks jev vs 30+ open models with 130K questions each
multimodalart · x · 2026-09-22
multimodalart introduced Decision Index 0.1, a rigorous leaderboard comparing jev with 30+ open-weight decision models: 35+ benchmarks, 130K questions per model, testing knowledge, automation, understanding and creativity.
Verdict: jev still leads open models with a margin in overall knowledge (probably a bigger model), but on tool use/automation, retrieval and classification, the gap is now close.
More from Models
- AI crowd discovers LLMs aren't always the cheapest, most effective tool — evilsocket · 2026-09-22
- cloneofsimo: academia badly underestimates the problems OpenAI's math agents are solving — cloneofsimo · 2026-09-22
- Founder finds asking the model to compare outputs restores drifting Astra quality — i_dg23 · 2026-09-22
- Community wonders if Alibaba has abandoned its Qwen 35B A3B small MoE line — Akainu_Fan · 2026-09-22
- Claude Opus 'acting like Sonnet' fuels speculation of new model launch — RyanMorrisonJer · 2026-09-22
- Reddit proposes measuring LLMs by cost per accepted task, not cost per token, after Grok 4.7 launch — Crescitaly · 2026-09-22