Decision Index 0.1: jev still leads open models, but tool use gap narrows fast
multimodalart · x · 2026-09-22
multimodalart released Decision Index 0.1, a rigorous leaderboard running 35+ benchmarks and 130K questions per model across knowledge, automation, understanding and creativity, comparing jev against 30+ open-weight decision models.
Findings: jev keeps the lead on overall knowledge (likely a larger model), but open models are now close on tool use/automation, retrieval and classification. Right behind: inference techniques enabling one-pass inference of Qwen3.8-27B and diffusion gemma without fine-tuning, and 3rd-place Decider 35B-A3B, a strong fine-tune of Qwen3.5 35B base.
More from Models
- AI crowd discovers LLMs aren't always the cheapest, most effective tool — evilsocket · 2026-09-22
- cloneofsimo: academia badly underestimates the problems OpenAI's math agents are solving — cloneofsimo · 2026-09-22
- Founder finds asking the model to compare outputs restores drifting Astra quality — i_dg23 · 2026-09-22
- Community wonders if Alibaba has abandoned its Qwen 35B A3B small MoE line — Akainu_Fan · 2026-09-22
- Claude Opus 'acting like Sonnet' fuels speculation of new model launch — RyanMorrisonJer · 2026-09-22
- Reddit proposes measuring LLMs by cost per accepted task, not cost per token, after Grok 4.7 launch — Crescitaly · 2026-09-22