llama.cpp adds decision models: typed questions, per-option probabilities in one forward pass
ngxson · x · 2026-10-02
llama.cpp server now supports decision models via the /v1/systemone endpoint: send a state (text, JSON, or screenshot) plus typed questions, and get a probability for every option in a single forward pass—no token-by-token generation or output parsing. The API follows the System One format from TypeSafe's Jev model, so existing clients only need a new base URL.
Supported open models (5, from 144M to 27B):
- Julia-1: 144M, mmBERT-small-based, 50+ languages, 3 ms per question
- Laya: 421M, ModernBERT-large-based, English, 5 ms
- Kev-4B / lev: 4B, Qwen3.5-based, 12/36 ms
- OpenJev: 27B, Qwen3.8-27B-based, multilingual (en/de/fr/hi/zh/ja), image input, 43 ms
Typical uses: request routing, content moderation, verifying an agent's step succeeded, or choosing an agent's next action. Question types include choice (options with per-option probabilities) and score. Quick start: llama serve -hf ggml-org/Kev-4B-GGUF. Models are collected on Hugging Face with a community Decision Index for comparisons.
Related event: llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models(3 posts)→
More from coding & agent
- Aviation's ASD-STE100 controlled language as an anti-AI-slop prompt hack, and where it fails — Paimaamu · 2026-10-02
- Pi Durable as statecharts: an interactive demo of crash-safe LLM agent harnesses — sloppenheimer · 2026-10-02
- Microsoft open-sources NVX, an ultra-light OpenVMM-based micro-VM sandbox for agentic workloads — unixterminal · 2026-10-02
- exe.dev's 'Run Fewer Agents': why task management isn't the fix for agent sprawl — charles_irl · 2026-10-02
- Building an agentic ML team: multi-agent pipeline with 40% token savings — kmeanskaran · 2026-10-02
- Claude Code creator: I don't prompt anymore, I write loops — a PM starter — aakashgupta · 2026-10-02