llama.cpp Adds Decision Model Endpoint with 3ms Inference for 144M Models

llama.cpp server now supports Decision Models via the new /v1/systemone endpoint, returning option probabilities in a single forward pass. It already works with five open-source models, with the smallest at 144M parameters and inference as fast as 3ms.

2026-10-02 ~ 2026-10-02 · 3 related posts

1 near-duplicate retellings: ggerganov