DeepSeek v4 and GLM Now Run Faster Than vLLM and SGLang
jedisct1 · x · 2026-09-09
Developer steeve reports that DeepSeek v4 and GLM now run faster on their inference stack than both vLLM and SGLang, the two dominant serving frameworks, adding that plenty of feature work remains but the performance milestone is reached. A notable third-party signal on next-gen open model inference performance.
More from Infra
- Baseten tops Coval's voice AI benchmark: STT ~5x faster than OpenAI with lowest WER — baseten · 2026-09-09
- 2×4090 llama.cpp concurrency: soft cap of 5 agents at 64k context, hard cap 9 — three weeks of data — Iamisseibelial · 2026-09-09
- exe.dev deep dive: ssh to a persistent Linux VM in half a second, priced like a folder — davidcrawshaw · 2026-09-09
- GPT-6 Astra lands on Amazon Bedrock with 1M-token context and first Critical cyber rating — AWS ML Blog · 2026-09-09
- Estha Turns One Mac Into a Shared Local AI Server for a Whole Team — HaktanSuren · 2026-09-09
- Alexandr Wang backs claim that Scale's own compute lets it subsidize Muse's speed — alexandr_wang · 2026-09-09