vLLM: agent sessions median 43 turns, 142K-token inputs vs 444-token outputs

vllm_project · x · 2026-09-09

vLLM profiled real coding-agent traffic: median 43 turns per session, 142K-token median input against only 444-token median output, 96%+ prefix-cache hit rate, and 44% of sessions forking subagents — long prefixes, tiny outputs, constant reuse. Companion data shows DeepSeek V4 Pro sustaining 83K tokens per GPU-second at a strict p90 interactivity SLO of 50+ tok/s per user (130K at the frontier tier), while serving the same workload on Opus 5 at the same cache-hit rate costs 106x more, highlighting the headroom open-weight models have with a tuned serving stack.

Related event: vLLM Unveils AgentX Benchmark: Up to 106x Cost Advantage for Real-World Agentic Serving(5 posts)→

Original post →

More from coding & agent

coding & agent channel →