llama.cpp adds /v1/systemone endpoint to run decision models locally, returning probabilities

solyarisoftware · x · 2026-10-03

A new llama.cpp update lets you run decision models locally: give it a question and choices, get probabilities back—no generated answer. The new /v1/systemone endpoint targets request routing, tool selection, and checking an agent's work.

Supported GGUFs: Julia-1 (144M), Laya, lev, Kev-4B, and OpenJev (27B, with image support). OpenJev's weights are CC BY-NC 4.0; the other four are Apache 2.0. Update llama.cpp to try it.

Related event: llama.cpp adds Decision Models: one forward pass scores options directly(8 posts)→

Original post →

More from Infra

Infra channel →