llama.cpp adds /v1/systemone endpoint to run decision models locally, returning probabilities
solyarisoftware · x · 2026-10-03
A new llama.cpp update lets you run decision models locally: give it a question and choices, get probabilities back—no generated answer. The new /v1/systemone endpoint targets request routing, tool selection, and checking an agent's work.
Supported GGUFs: Julia-1 (144M), Laya, lev, Kev-4B, and OpenJev (27B, with image support). OpenJev's weights are CC BY-NC 4.0; the other four are Apache 2.0. Update llama.cpp to try it.
Related event: llama.cpp adds Decision Models: one forward pass scores options directly(8 posts)→
More from Infra
- SemiAnalysis: Nvidia's custom NVHBM frees ~25% more compute die area on Feynman — zephyr_z9 · 2026-10-03
- gufo-Qwen3.6-35B hits 3095 tok/s prefill, 190 tok/s decode on Strix Halo — nubela · 2026-10-03
- Garage server farms return: GPU and power shortages reverse the AWS era in Palo Alto — bookwormengr · 2026-10-03
- AI maxi calls GPU price hike a bubble peak, plans to buy cheap cards after burst — AIFlow_ML · 2026-10-03
- SMIC Roadmap Reportedly Shows No EUV Production by 2030, TSMC Gap May Exceed 10 Years — teortaxesTex · 2026-10-03
- Microsoft sends controversial Majorana topological qubit chips to DARPA for independent testing — skdh · 2026-10-03