llama.cpp adds decision models: score options in one forward pass, from 144M Julia-1 at 3ms
ggerganov · x · 2026-10-02
llama.cpp now supports decision models via a new /v1/systemone endpoint, following the System One format from TypeSafe's Jev model.
Unlike chat models that generate tokens one at a time, a decision model reads a state (text, JSON, or a screenshot) once and returns a probability for each option you provide in a single forward pass. Typical uses include request routing, content moderation, verifying agent steps, and choosing an agent's next action.
Supported models:
- Julia-1 (144M, mmBERT-small based, 50+ languages, Apache 2.0, median 3ms per question)
- Laya (421M, ModernBERT-large, English, 5ms)
- Kev-4B (4B, Qwen3.5-4B-Base, 12ms)
- lev (4B, Qwen3.5-4B, 36ms)
- OpenJev (27B, multilingual incl. Chinese, image-capable, CC BY-NC 4.0, 43ms)
Benchmarks were run on one NVIDIA RTX PRO 6000. The API supports choice and score question types. Quick start: llama serve -hf ggml-org/Kev-4B-GGUF.
Related event: llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models(3 posts)→
More from coding & agent
- Intent launches Stacks UI for seeing and steering fleets of AI agents — LukeW · 2026-10-03
- WebMCP lets pages expose tools to AI agents directly, replacing screenshot-based clicking — TejasKumar_ · 2026-10-03
- City Farmers joins OpenAI's first Codex Physical Builds batch with a Raspberry Pi setup — broodsugar · 2026-10-02
- SWE-sweep benchmark tests whether AI agents can find bugs before users hit them — klieret · 2026-10-02
- Turn Claude's Dot into a live DJ with browser use and your calendar — pvncher · 2026-10-02
- Indie dev builds 'AdSense for AI agents' with free search and 70% ad revenue split — Keats0206 · 2026-10-02