Jev benchmark results: strong at tool calling and automation, weak at reasoning and retrieval
multimodalart · x · 2026-09-22
multimodalart shares how the model Jev fares on a deliberately challenging benchmark designed to avoid instant saturation.
Highlights:
- Particularly strong at tools and automation tasks
- Handles knowledge-based questions up to graduate level, Python function choosing, tool calling, and model routing
Weaknesses: still struggles with puzzle and reasoning games like chess, does poorly on super-hard document retrieval tasks that normally require reasoning models, and underperforms on recognizing chords from musical notes and fine-grained sentiment analysis.
Related event: Jev Model Shines in Tool Use but Lags in Calibration(4 posts)→
More from Models
- Claude counts tokens, not messages: 9 tricks to avoid hitting usage limits — HeyAmit_ · 2026-09-22
- Leaked screenshots surface of rumored OpenAI "Aeon" persistent agent — PrisonOfH0pe · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Pelican SVG test puts unreleased GPT-6 Astra head-to-head with Anthropic's Mythos — PrisonOfH0pe · 2026-09-22
- Gemini Pro users report web access silently disabled, even on paid subscription — zsolt67 · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22