Decision Index 0.1: leaderboard asks 130K questions to 30+ open decision models
multimodalart · x · 2026-09-22
multimodalart launched Decision Index 0.1, a rigorous leaderboard for decision models.
- Compares Jev against 30+ open-weight decision models
- Spans 35+ benchmarks, asking 130K questions to each model
- Tests knowledge, automation, understanding, and even creativity
A rare large-scale evaluation focused on decision-making rather than chat ability.
More from Models
- Claude counts tokens, not messages: 9 tricks to avoid hitting usage limits — HeyAmit_ · 2026-09-22
- Leaked screenshots surface of rumored OpenAI "Aeon" persistent agent — PrisonOfH0pe · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Pelican SVG test puts unreleased GPT-6 Astra head-to-head with Anthropic's Mythos — PrisonOfH0pe · 2026-09-22
- Gemini Pro users report web access silently disabled, even on paid subscription — zsolt67 · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22