llama.cpp adds decision models: /v1/systemone scores options in a single forward pass
ggerganov · x · 2026-10-02
llama.cpp server now ships a /v1/systemone endpoint for decision models: you send a state (text, JSON, or a screenshot) plus typed questions, and the model returns a probability for each option in a single forward pass—no token-by-token generation or output parsing. The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL (implementation in PR #29818).
Typical uses: request routing, content moderation, verifying an agent's step succeeded, or picking an agent's next action.
First models (median latency per question on one RTX PRO 6000):
- Julia-1: 144M, based on mmBERT-small, 50+ languages, Apache 2.0, 3ms
- Laya: 421M, ModernBERT-large, English, Apache 2.0, 5ms
- Kev-4B: 4B, Qwen3.5-4B-Base, English, Apache 2.0, 12ms
- lev: 4B, Qwen3.5-4B, English, Apache 2.0, 36ms
- OpenJev: 27B, Qwen3.8-27B, en/de/fr/hi/zh/ja, image support, CC BY-NC 4.0, 43ms
Quick start: llama serve -hf ggml-org/Kev-4B-GGUF. Questions include choice (returns top option plus per-option probabilities) and score types. Models live in the Hugging Face Decision models collection, with a community Decision Index for comparisons and more models on the way.
Related event: llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models(3 posts)→
More from coding & agent
- Aviation's ASD-STE100 controlled language as an anti-AI-slop prompt hack, and where it fails — Paimaamu · 2026-10-02
- Pi Durable as statecharts: an interactive demo of crash-safe LLM agent harnesses — sloppenheimer · 2026-10-02
- Microsoft open-sources NVX, an ultra-light OpenVMM-based micro-VM sandbox for agentic workloads — unixterminal · 2026-10-02
- exe.dev's 'Run Fewer Agents': why task management isn't the fix for agent sprawl — charles_irl · 2026-10-02
- Building an agentic ML team: multi-agent pipeline with 40% token savings — kmeanskaran · 2026-10-02
- Claude Code creator: I don't prompt anymore, I write loops — a PM starter — aakashgupta · 2026-10-02