vLLM releases open Decision 2.0 models: 64 questions in one pass at 63ms

vllm_project · x · 2026-10-03

The vLLM Semantic Router team released Decision 2.0, Apache-2.0 open decision models spanning 0.6B to 27B that load with Transformers. The 0.6B/0.8B/2B/4B variants rank #1 at their size on the Jev Decision Index (27B is #3 overall), scoring up to 2.5× Decision 1.0 at equal size. One forward pass answers 64 questions about a request in 63ms on a single GPU—roughly 20× faster than the Jev API per decision. Built by Xunzhuo Liu's team for vLLM Semantic Router.

Original post →

More from Models

Models channel →