vLLM hits 130k tok/s on DeepSeek V4 Pro in AgentX benchmark
AccBalanced · x · 2026-08-27
SemiAnalysis released AgentX 1.0, an open-source multi-turn agentic coding benchmark from $3M of real traces. vLLM showed competitive performance: DeepSeek V4 Pro reached 130,093 tok/s/chip, MiniMax M3 77,079 tok/s/chip, and Kimi K3 12,479 tok/s/chip. This highlights vLLM's performance on real-world agentic workloads.
More from Infra
- Nvidia guides $108B Q3 revenue, doubling growth even without China data center sales — inductionheads · 2026-08-27
- Palantir Karp: Serious enterprises need to own their models and infrastructure — JosephJacks_ · 2026-08-27
- Economics of Becoming an OpenRouter Provider with H200 Nodes — ell-hol1 · 2026-08-27
- Analysis layers exist between Nvidia sales and enterprise ROI — iamKierraD · 2026-08-27
- Apple still leads in laptop processor performance years after M1, leaving Intel and AMD behind — lemire · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27