Decision Index 0.1 Benchmarks 30+ Open Models with 132K Decisions Each
multimodalart (a Hugging Face team member) has released Decision Index 0.1, a rigorous leaderboard for open-weight decision models. The project uses 35+ benchmarks and asks each model 130,000 questions (victormustar cites the exact figures as 132,422 decisions across 37 benchmarks), with no truncation, comparing Jev against 30+ open-weight decision models across four dimensions: knowledge, automation, comprehension, and creativity. The conclusion: Jev still leads open-source models at 59.5 points, and the gap in tool-use tasks is narrowing.
Confirmed
- Decision Index 0.1 was released by multimodalart on 09-22 as a Hugging Face team project
- Evaluation scale: 35+ benchmarks (victormustar mentions 37), 132,422 decisions, roughly 130,000 questions per model
- Evaluated models were Jev and its reproductions, totaling 30+ open-weight decision models
- Evaluation dimensions include knowledge, automation, comprehension, and creativity
- Jev ranks first at 59.5 points and remains the open-source leader
Why it matters
- According to victormustar, Jev already had 31 reproduction versions within a week of release, but most prior comparisons only ran a few hundred decisions — small samples and unreliable conclusions
- With six-figure decision testing and no truncation, Decision Index 0.1 offers a statistically more credible cross-model benchmark for decision models
2026-09-22 ~ 2026-09-22 · 7 related posts
Primary sources
- Decision Index 0.1 benchmarks 30+ open-weight decision models across 35+ evals and 130K questions — multimodalart · 2026-09-22
- Decision Index 0.1 benchmarks jev vs 30+ open models with 130K questions each — multimodalart · 2026-09-22
- 37 benchmarks, 130K decisions per model: local LLMs tested on one RTX 6000 PRO — multimodalart · 2026-09-22
- 37 benchmarks, 130K decisions per model: jev excels at tools and automation — multimodalart · 2026-09-22
- [source] New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
2 near-duplicate retellings: multimodalart · multimodalart