Ollama adds decision models: Nimble 9B runs locally with ~91ms decisions
ollama · x · 2026-09-30
Ollama's v0.35 introduces a /v1/systemone endpoint supporting TypeSafe's Jev-style decision models, with three new models available for local inference:
- nimble: an open-source 9B decision model from Bespoke Labs, fine-tuned from Qwen3.5-9B under Apache 2.0. It takes text plus a set of named questions and returns typed answers with probabilities in a single call—no reasoning step, it scores answer tokens directly. Up to 64 questions per request, averaging 91ms per decision on an M5 Max MacBook Pro.
- tev1 4b / 0.8b: experimental decision models from Together AI.
Use cases include ticket triage, model routing, and content/safety moderation. Bespoke Labs benchmarks across 13 public datasets (3,880 human-labeled decisions) show Nimble and Tev1 on Ollama approaching Jev 1.13 accuracy. Nimble is trained on contrastive pairs differing by one fact that flips the answer.
Related event: Ollama 0.35 Adds Local Decision Models, Nimble Responds in 91ms(3 posts)→
More from coding & agent
- Trust is not a security primitive: why AI agent sandboxes must assume nothing about the model — basedjensen · 2026-09-30
- Omni-IO Skills: harness-level composition lifts GPT-5.6's omni-input support from 40% to 100% — NationalUniversityofSingapore · 2026-09-30
- Omni-Decision: evidence-ledger planning hits 81.4% on OmniGAIA at 43% of Gemini-3.1-Pro's cost — Ming Ma · 2026-09-30
- TypeSafe's official Claude Code skill: batching 13 questions in one call is 12.2x cheaper and 10x faster — Roger_M_Taylor · 2026-09-30
- AgentCraft: StarCraft rebuilt in the browser as an arena for AI agents — philipvollet · 2026-09-30
- Developer says AI agents did $100K worth of human labor in under a day — RileyRalmuto · 2026-09-30