DeepSeek v4 and GLM Now Run Faster Than vLLM and SGLang

jedisct1 · x · 2026-09-09

Developer steeve reports that DeepSeek v4 and GLM now run faster on their inference stack than both vLLM and SGLang, the two dominant serving frameworks, adding that plenty of feature work remains but the performance milestone is reached. A notable third-party signal on next-gen open model inference performance.

Original post →

More from Infra

Infra channel →