llama.cpp adds decision models: /v1/systemone scores options in a single forward pass

ggerganov · x · 2026-10-02

llama.cpp server now ships a /v1/systemone endpoint for decision models: you send a state (text, JSON, or a screenshot) plus typed questions, and the model returns a probability for each option in a single forward pass—no token-by-token generation or output parsing. The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL (implementation in PR #29818).

Typical uses: request routing, content moderation, verifying an agent's step succeeded, or picking an agent's next action.

First models (median latency per question on one RTX PRO 6000):

Quick start: llama serve -hf ggml-org/Kev-4B-GGUF. Questions include choice (returns top option plus per-option probabilities) and score types. Models live in the Hugging Face Decision models collection, with a community Decision Index for comparisons and more models on the way.

Related event: llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →