Jev excels at graduate-level knowledge, Python function choosing, tool calling and model routing
multimodalart · x · 2026-09-22
multimodalart adds more Jev benchmark details: the model performs particularly well on knowledge-based questions up to graduate level, Python function choosing, tool calling, and model routing. This complements earlier findings that Jev struggles with reasoning games like chess, hard document retrieval, chord recognition, and fine-grained sentiment analysis.
Related event: Jev Model Shines in Tool Use but Lags in Calibration(4 posts)→
More from Models
- Claude counts tokens, not messages: 9 tricks to avoid hitting usage limits — HeyAmit_ · 2026-09-22
- Leaked screenshots surface of rumored OpenAI "Aeon" persistent agent — PrisonOfH0pe · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Pelican SVG test puts unreleased GPT-6 Astra head-to-head with Anthropic's Mythos — PrisonOfH0pe · 2026-09-22
- Gemini Pro users report web access silently disabled, even on paid subscription — zsolt67 · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22