vLLM releases open Decision 2.0 models: 64 questions in one pass at 63ms
vllm_project · x · 2026-10-03
The vLLM Semantic Router team released Decision 2.0, Apache-2.0 open decision models spanning 0.6B to 27B that load with Transformers. The 0.6B/0.8B/2B/4B variants rank #1 at their size on the Jev Decision Index (27B is #3 overall), scoring up to 2.5× Decision 1.0 at equal size. One forward pass answers 64 questions about a request in 63ms on a single GPU—roughly 20× faster than the Jev API per decision. Built by Xunzhuo Liu's team for vLLM Semantic Router.
More from Models
- Steve Yegge: Two Weeks With Opus 5.5 — Precision Rivals Fable, Recall Trails on Open-Ended Tasks — Steve_Yegge · 2026-10-03
- OpenAI's Codex global reset appears to skip Business accounts, support suggests buying credits — AdventurousFeeling19 · 2026-10-03
- Mystery 'iguana_necktie' Field Spotted in Anthropic Usage API — bytebot · 2026-10-03
- Insider teases 'new SSI model,' calling Ilya 'a truly remarkable human being' — iruletheworldmo · 2026-10-03
- Exploit Bench results are highly harness-sensitive: GLM 5.3 Flash beats 4.1, best in Claude Code — teortaxesTex · 2026-10-03
- Reddit user says Opus 5.5 burns weekly cap at 300M tokens, down from 1-2B before — Chemical-Ad-7982 · 2026-10-03