llama.cpp adds Decision Models: one forward pass scores options directly
ggerganov, author of llama.cpp, announced that the latest build adds support for "Decision Models" via a new server endpoint /v1/systemone: given a state (text, JSON, or a screenshot) and several typed questions, the model returns a probability for each option directly in a single forward pass, instead of generating text token by token and then parsing it. The feature has drawn wide attention in the community.
Confirmed
- Released by ggerganov himself; the endpoint is /v1/systemone, following the System One format introduced by the TypeSafe Jev model, and the API is compatible with that format.
- Input states can be text, JSON, or screenshots, and questions carry type annotations.
- The smallest model has only 144M parameters, with inference as fast as roughly 3ms; ngxson confirmed 5 open-source models are already supported, with more on the way.
- Existing clients can use it by simply changing the base URL, enabling local, efficient, and private Jev-style inference.
- The official ggml blog on Hugging Face has published an introductory post (decision-models-in-llamacpp).
Why it matters
- Decision models turn "generative" calls into "scoring" calls, eliminating the overhead of token-by-token generation and the uncertainty of output parsing—well suited to high-frequency, structured decision scenarios.
- With 3ms-level local inference, lightweight models become viable for real-time interaction and privacy-sensitive use cases, opening up a new way to play with local inference.
- Prominent developers such as mitsuhiko (Flask author Armin Ronacher) reshared it, underscoring the feature's influence in the developer community.
2026-10-02 ~ 2026-10-03 · 8 related posts
Primary sources
- llama.cpp adds decision models: /v1/systemone scores options in a single forward pass — ggerganov ·
- llama.cpp adds decision models: score options in one forward pass, from 144M Julia-1 at 3ms — ggerganov ·
- llama.cpp adds decision models: typed questions, per-option probabilities in one forward pass — ngxson ·
- llama.cpp adds Decision Models, expanding local inference capabilities — paf1138 · 2026-10-02
- [source] llama.cpp adds decision models: typed questions, per-option probabilities in one forward pass — ngxson · 2026-10-02
- llama.cpp ships /v1/systemone endpoint for local Jev-style decision models — mitsuhiko · 2026-10-03
- llama.cpp adds /v1/systemone endpoint to run decision models locally, returning probabilities — solyarisoftware · 2026-10-03
4 near-duplicate retellings: ggerganov · ggerganov · adnan_hashmi · kalyan_kpl