FULL STORY
Decision Index: Open Decision Model Leaderboard Iterates Fast
HuggingFace's multimodalart launched Decision Index 0.1 on Sept 22, ranking 30+ open decision models over 130k decisions. Two days later, v0.2 improved the formula, added 50 models, and saw Perplexity's CTO model top the board.
2026-09-22 ~ 2026-09-24 · 2 episodes · 9 posts
Episode 1 · Decision Index 0.1 Benchmarks 30+ Open-Weight Decision Models with 132k Decisions, Jev Leads (2026-09-22, 7 posts)
multimodalart (Hugging Face team) released Decision Index 0.1 on 09-22, a rigorous leaderboard for open-weight decision models. It covers 37 benchmarks across 5 task categories, runs 132,422 decisions (130k questions per model) without truncation, and compares Jev and its reproductions among 30+ open-weight decision models across knowledge, automation, comprehension and creativity dimensions. Jev leads open models at 59.5, and the tool-use gap is narrowing. All tests ran on a single RTX 6000 Pro, with benchmarks deliberately kept hard to avoid immediate saturation.
Confirmed
- Released by multimodalart on 09-22 as a Hugging Face team project
- Scale: 37 benchmarks, 5 categories, 132,422 decisions, 130k questions per model (some posts by multimodalart say 35+ benchmarks; victormustar cites 37)
- Evaluated Jev and its reproductions, 30+ open-weight decision models total
- Dimensions: knowledge, automation, comprehension, creativity
- Jev ranks first at 59.5; tool-use gap narrowing; top of the leaderboard also mentions Qwen3.8-27B and diffusion gemma single-inference entries
- All runs on a single RTX 6000 Pro; benchmarks deliberately difficult to avoid saturation on release
Why it matters
- Per victormustar, Jev already has 31 reproductions within a week of release, but most prior comparisons ran only a few hundred decisions — too small to be reliable
- With six-figure decision counts, no truncation and unified hardware, Decision Index 0.1 offers a statistically more credible benchmark for decision models
- Decision Index 0.1 benchmarks 30+ open-weight decision models across 35+ evals and 130K questions — multimodalart · 2026-09-22
- Decision Index 0.1: leaderboard asks 130K questions to 30+ open decision models — multimodalart · 2026-09-22
- Decision Index 0.1 benchmarks jev vs 30+ open models with 130K questions each — multimodalart · 2026-09-22
- Decision Index 0.1: jev still leads open models, but tool use gap narrows fast — multimodalart · 2026-09-22
- 37 benchmarks, 130K decisions per model: local LLMs tested on one RTX 6000 PRO — multimodalart · 2026-09-22
- 37 benchmarks, 130K decisions per model: jev excels at tools and automation — multimodalart · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
Episode 2 · Decision Index 0.2 released; Perplexity CTO's model tops open ranking (2026-09-24, 2 posts)
Multimodal Art released Decision Index 0.2 with an improved formula, 29 new Jev models and 21 benchmarks. Perplexity's CTO topped the open-source ranking with AutoJev-27B, trained for $3,000.
- Decision Index 0.2 Adds 29 Models and 21 Benchmarks; AutoJev-27B Takes Open-Weights Lead — multimodalart · 2026-09-24
- Perplexity CTO's AutoJev-27B Takes Open-Source Lead on Decision Index, Trained for Just $3k — antoine_chaffin · 2026-09-24