Open-weight models close in on TypeSafe's Jev: Shisa DE-1 just 6 points behind, wins on some tasks
drdanielbender · x · 2026-09-28
jtdavies recreated Rogerio's LangWatch benchmark (11 tasks) on an M5 Pro Mac mini and ran 15 open-weight models against TypeSafe's decision model Jev:
- Best was Shisa DE-1, 6 points behind Jev's published average, level or ahead on several tasks
- 4B-class models trailed by only a few points, running locally at 120 ms per decision
- On Matthew Berman's 9 use cases, Shisa vs a local 4B reproduction of Jev's approach was nearly identical (5 ties, 2 wins each) at a bit over half the speed
Under two weeks after Jev's launch, open weights are already nearly level — strong evidence for small models on decision tasks.
More from Infra
- OriginTrail ships DKG V10.0.19 on mainnet for faster AI agent context graphs — melnykowycz · 2026-09-29
- INT21 claims 20 AI-generated inference engines in 2 weeks, MiMo hits 1,308 tok/s — bingxu_ · 2026-09-29
- Agent-driven synthetic monitoring with Amazon Nova Act replaces brittle UI scripts — AWS ML Blog · 2026-09-28
- NVIDIA demos 2,529 output tokens per second for Qwen 27B at GTC — IanAndrewsDC · 2026-09-28
- Automating Amazon Textract adapter lifecycle management across accounts — AWS ML Blog · 2026-09-28
- Muse is one CPU-heavy implementation, not a proxy for sizing the agentic CPU opportunity — BenBajarin · 2026-09-28