37 benchmarks, 130K decisions per model: jev excels at tools and automation

multimodalart · x · 2026-09-22

multimodalart details a custom evaluation: 37 benchmarks across 5 task categories, 130K decisions per model, all run on identical hardware (1x RTX 6000 PRO). The benchmark is deliberately hard to avoid instant saturation. Results show jev performing especially well on tools and automation tasks.

Related event: Decision Index 0.1 Launches: 132k Decisions Benchmark Jev Against 30+ Open Models(7 posts)→

Original post →

More from Models

Models channel →