Decision Index 0.1 Benchmarks 30+ Open Models with 132K Decisions Each

multimodalart (a Hugging Face team member) has released Decision Index 0.1, a rigorous leaderboard for open-weight decision models. The project uses 35+ benchmarks and asks each model 130,000 questions (victormustar cites the exact figures as 132,422 decisions across 37 benchmarks), with no truncation, comparing Jev against 30+ open-weight decision models across four dimensions: knowledge, automation, comprehension, and creativity. The conclusion: Jev still leads open-source models at 59.5 points, and the gap in tool-use tasks is narrowing.

Confirmed

Why it matters

2026-09-22 ~ 2026-09-22 · 7 related posts

Primary sources

2 near-duplicate retellings: multimodalart · multimodalart