vLLM hits 130k tok/s on DeepSeek V4 Pro in AgentX benchmark

AccBalanced · x · 2026-08-27

SemiAnalysis released AgentX 1.0, an open-source multi-turn agentic coding benchmark from $3M of real traces. vLLM showed competitive performance: DeepSeek V4 Pro reached 130,093 tok/s/chip, MiniMax M3 77,079 tok/s/chip, and Kimi K3 12,479 tok/s/chip. This highlights vLLM's performance on real-world agentic workloads.

Original post →

More from Infra

Infra channel →