vLLM: agent sessions median 43 turns, 142K-token inputs vs 444-token outputs
vllm_project · x · 2026-09-09
vLLM profiled real coding-agent traffic: median 43 turns per session, 142K-token median input against only 444-token median output, 96%+ prefix-cache hit rate, and 44% of sessions forking subagents — long prefixes, tiny outputs, constant reuse. Companion data shows DeepSeek V4 Pro sustaining 83K tokens per GPU-second at a strict p90 interactivity SLO of 50+ tok/s per user (130K at the frontier tier), while serving the same workload on Opus 5 at the same cache-hit rate costs 106x more, highlighting the headroom open-weight models have with a tuned serving stack.
More from coding & agent
- Anthropic's Claude Tag acts as on-call first responder for CI/CD failures — ClaudeDevs · 2026-09-09
- Agents now ship with a soul.md file, a nod to Peter Steinberger's influence — altryne · 2026-09-09
- How do you coordinate 10,000 AI agents on one problem? Hierarchical orchestration ideas emerge — pwlot · 2026-09-09
- Mitchell Hashimoto demos Superlogical remote persistent sessions, a full SSH replacement — iannuttall · 2026-09-09
- Catching LLM page overflow: a measure-and-repair loop for single-page LaTeX generation — Scholeristical · 2026-09-09
- OpenClaw ships 2026.9.3 with live browser automation, revocable share links, cloud repo work — heyneighbor · 2026-09-09