DeepSeek Uses More Tokens but is Faster per Single Request

Sentdex · x · 2026-07-03

Sentdex observed that DeepSeek consumes about 2.5x more tokens than GLM 5.2 IQ4 but is 5x faster at batch size=1. Once concurrency is introduced, vLLM's concurrent processing capabilities take off rapidly. He also noticed a performance drop in vLLM tensor parallelism at batch size=2 before it recovers, and asked the community for insights on this behavior.

Original post →

More from Infra

Infra channel →