llama.cpp adds decision models: typed questions, per-option probabilities in one forward pass

ngxson · x · 2026-10-02

llama.cpp server now supports decision models via the /v1/systemone endpoint: send a state (text, JSON, or screenshot) plus typed questions, and get a probability for every option in a single forward pass—no token-by-token generation or output parsing. The API follows the System One format from TypeSafe's Jev model, so existing clients only need a new base URL.

Supported open models (5, from 144M to 27B):

Typical uses: request routing, content moderation, verifying an agent's step succeeded, or choosing an agent's next action. Question types include choice (options with per-option probabilities) and score. Quick start: llama serve -hf ggml-org/Kev-4B-GGUF. Models are collected on Hugging Face with a community Decision Index for comparisons.

Related event: llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →