Simulated leaderboard: Jev Router posts +3.6% net gain at $0.155 per task, near Opus steerability
arena · x · 2026-10-08
A follow-up from Agent Arena's Jev Router evaluation presenting its simulated leaderboard results:
- +3.6% net improvement at a median cost of $0.155 per task.
- It would rank below DeepSeek V4.1 Flash, the model it routes to most, but beat it on Steerability (+10.0% vs -0.2%) and Praise vs Complaint (+6.3% vs +2.1%).
- DeepSeek leads on Confirmed Success (+8.3% vs +2.2%), Bash Recovery (+7.8% vs -0.7%), and Tool Hallucination (+0.4% vs +0.2%).
- The key signal remains steerability: only 0.48 percentage points below Claude Opus 5.5 (High)'s +10.48%.
More from Models
- Skeptical of 'small model for evals' moats: labs will distill it into cheaper models — Shahules786 · 2026-10-08
- Bengio disputes 'just a sandbox bug' framing of AI agent hacks in FT op-ed — AlexTensor · 2026-10-08
- OpenAI model proves Hilbert's Tenth Problem false over Q, sidestepping 80-year approach — aran_nayebi · 2026-10-08
- What Anthropic's $200 tier changes about choosing between Opus and Sonnet — thursdai_pod · 2026-10-08
- AI flip: it may plan your Boston trip before solving the Riemann hypothesis — jxmnop · 2026-10-08
- OpenRouter's Usage and Spend Charts Have Split: Used Models Aren't the Paid Ones — Spiritual_Skirt_9312 · 2026-10-08