OSDI Paper: vLLM and Others Unsuitable for Long-Running Coding Agents
AccBalanced · x · 2026-07-19
At USENIX OSDI, researchers from UW and UVA presented a paper on tuning vLLM for coding agents. - **Background**: They collected the most comprehensive Claude Code OS-level tracing dataset from their student labs. - **Core Finding**: Existing serving frameworks like vLLM and SGLang cannot directly handle long-running agent workloads. - **Technical Detail**: If KV cache is constrained, more CPU scheduling overhead is needed to predict cache reuse in long coding sessions.
More from coding & agent
- Codex Computer Use grabs a voicemail access code through iPhone call screening — KarelDoostrlnck · 2026-07-21
- Grok Imagine’s Agent mode adds image cropping for smoother video transitions — elonmusk · 2026-07-21
- An engineering platform that auto-assigns and resolves Jira bugs with AI agents — Pavan_Belagatti · 2026-07-21
- Reddit asks how to stop agents from retrying the same failed call until budgets burn — bulleykebaal · 2026-07-21
- A “shadow steering” pattern lets a larger model intervene only when a smaller one starts to fail — dotey · 2026-07-21
- Clanker Cloud teases easier agent-to-agent communication for its next update — tekbog · 2026-07-21