Decision Index 0.1 benchmarks 30+ open-weight decision models across 35+ evals and 130K questions
multimodalart · x · 2026-09-22
multimodalart released Decision Index 0.1, a rigorous evaluation of JEV and 30+ open-weight decision/classification models. The project runs each model through 35+ benchmarks with 130K questions per model, testing knowledge, automation, understanding and even creativity. Full results available at the linked site.
Related event: Decision Index 0.1 Benchmarks 30+ Open Decision Models with 130K Queries(4 posts)→
More from Models
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22
- Open-Source Perf-Latency Pareto Frontier Features Qwen3-4B and DiffusionGemma — multimodalart · 2026-09-22
- Reddit thread rounds up rumored models: Gemini 4.0, OpenAI 'Bel', K4 race to ASI — IllCryptographer9461 · 2026-09-22
- Reddit asks: why optimize local LLMs for coding, not world knowledge via N-gram? — ironicstatistic · 2026-09-22
- Tiny model frontier mapped: GLiNER 2.5 leads <350M, Kev family tops two brackets — multimodalart · 2026-09-22
- Jev service suspends signups after overflowing; user to push another billion tokens today — MaziyarPanahi · 2026-09-22