Decision Index 0.1 benchmarks jev vs 30+ open models with 130K questions each

multimodalart · x · 2026-09-22

multimodalart introduced Decision Index 0.1, a rigorous leaderboard comparing jev with 30+ open-weight decision models: 35+ benchmarks, 130K questions per model, testing knowledge, automation, understanding and creativity.

Verdict: jev still leads open models with a margin in overall knowledge (probably a bigger model), but on tool use/automation, retrieval and classification, the gap is now close.

Related event: Decision Index 0.1 Launches: 132K Decisions Evaluated, Jev Leads Open-Weight Models at 59.5(7 posts)→

Original post →

More from Models

Models channel →