37 benchmarks, 130K decisions per model: jev excels at tools and automation
multimodalart · x · 2026-09-22
multimodalart details a custom evaluation: 37 benchmarks across 5 task categories, 130K decisions per model, all run on identical hardware (1x RTX 6000 PRO). The benchmark is deliberately hard to avoid instant saturation. Results show jev performing especially well on tools and automation tasks.
More from Models
- Claude counts tokens, not messages: 9 tricks to avoid hitting usage limits — HeyAmit_ · 2026-09-22
- Leaked screenshots surface of rumored OpenAI "Aeon" persistent agent — PrisonOfH0pe · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Pelican SVG test puts unreleased GPT-6 Astra head-to-head with Anthropic's Mythos — PrisonOfH0pe · 2026-09-22
- Gemini Pro users report web access silently disabled, even on paid subscription — zsolt67 · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22