Ollama adds local Jev-style decision models; Nimble 9B makes calls in 91ms on M5 Max
mchiang0610 · x · 2026-09-30
Ollama 0.35 adds support for Jev-style decision models via a new /v1/systemone endpoint, answering a set of named questions in one local request.
- Use cases: ticket triage, model routing, content/safety moderation — tasks needing near-instant decisions.
- Performance: running locally eliminates network latency; Nimble 9B averaged 91ms per decision on an M5 Max while playing a Pac-Man demo in real time.
- Models available: nimble (open-source 9B from Bespoke Labs), plus experimental tev1 4B and 0.8B decision models from Together AI, with more coming including cloud-served ones.
- Benchmarks: Bespoke Labs' public suite (13 human-labeled datasets, 3,880 decisions) shows local Nimble/Tev1 accuracy on par with cloud Jev 1.13.
- Get started with ollama pull nimble.
Related event: Ollama 0.35 Adds Local Decision Models, Nimble Responds in 91ms(3 posts)→
More from Models
- DeepSeek's forgotten male persona: why the 'big fat fish' meme won the Bilibili war — teortaxesTex · 2026-09-30
- GPT-6.1 Sol appears silently nerfed mid-task, user reports 3B-level output quality — Ferzelibey · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- DeepSeek is giving users 6 yuan in free API credits via its harness — teortaxesTex · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30