Jev benchmark results: strong at tool calling and automation, weak at reasoning and retrieval

multimodalart · x · 2026-09-22

multimodalart shares how the model Jev fares on a deliberately challenging benchmark designed to avoid instant saturation.

Highlights:

Weaknesses: still struggles with puzzle and reasoning games like chess, does poorly on super-hard document retrieval tasks that normally require reasoning models, and underperforms on recognizing chords from musical notes and fine-grained sentiment analysis.

Related event: Jev Model Shines in Tool Use but Lags in Calibration(4 posts)→

Original post →

More from Models

Models channel →