DecideBench: Imajev-4B Hits 95% But No Open Jev Alternative Beats Jev Yet
choyiny · reddit · 2026-09-30
choyiny benchmarked open-source "typed decision" models (DecideBench) on a production-like dataset. Conclusion: no open model beats Jev yet, but some come close at lower cost.
Key results
- Imajev-4B: 95% accuracy — the cheapest model above 90%
- Tev (Together AI): 92.8%
- Decider-4B and JevK5: both >88%, worth testing at lower price/latency
- Kev4B, Laya, CLM, Julia-1 still struggle on the production-like dataset
The author invites suggestions for more Jev-like models to evaluate.
More from Models
- GPT-6.1 Sol hands-on: efficient but slow, and subscription changes worry users — kimmonismus · 2026-09-30
- GPT 6.1 Astra ultra code mode builds a three.js ocean scene — OpenAIDevs · 2026-09-30
- Analysis: OpenAI's DOTS may be the answer to ChatGPT's growing product sprawl — mark_k · 2026-09-30
- NaceAI's Drex 1.5, a sub-10B decision model, tops Decision Index among 71 entries — nischay_twt · 2026-09-30
- OpenAI researcher says GPT-6.1 Sol almost shipped with a much cuter name — keyanzhang · 2026-09-30
- Meta Muse Spark 1.3 takes #2 on ProgramBench at rock-bottom cost — jyangballin · 2026-09-30