Agent Arena: Jev Router costs 38% more than DeepSeek V4.1 Flash but matches Opus 5.5 steerability
arena · x · 2026-10-08
Agent Arena evaluated @typesafeai's Jev Router across 4,700+ real-world agentic sessions:
- No Pareto improvement: matching DeepSeek V4.1 Flash (Max) performance costs 38% more, with 1.7x higher median request latency (6.18s vs 3.64s; P90 gap widens to 30.59s vs 14.26s).
- Routing mix: 28.1% of calls went to DeepSeek V4.1 Flash, with a clear OpenAI preference — GPT-6 Astra (23.5%), Sol (22.8%), Luna (11.1%).
- If leaderboarded: +3.6% net improvement at $0.155 median cost per task, ranking below DeepSeek V4.1 Flash but winning on Steerability (+10.0% vs -0.2%) and Praise vs Complaint (+6.3% vs +2.1%), while trailing on Confirmed Success, Bash Recovery, and Tool Hallucination.
- Key strength is steerability: its +10% score sits just 0.48 points below Claude Opus 5.5 (High), showing it can route to stronger LLMs in response to user corrections.
More from coding & agent
- Durable Actors: open-source Durable Objects alternative built for agent fleets — ritakozlov · 2026-10-08
- Essential tech skills in the AI era: Linux, Claude Code, Git, LaTeX and more — jessi_cata · 2026-10-08
- Prompting is becoming more like hiring a human: make the model prove it understands your repo first — seanmcdonaldxyz · 2026-10-08
- Sentry CEO: benchmarks show nearly all code review models surface the same primary issues — zeeg · 2026-10-08
- Meta, UW et al. open-source Context Language Models that natively manage their own context — RulinShao · 2026-10-08
- Dev says he can't remember the last time he ran a git command by hand — haydendevs · 2026-10-08