OSDI Paper: vLLM and Others Unsuitable for Long-Running Coding Agents
AccBalanced · x · 2026-07-19
At USENIX OSDI, researchers from UW and UVA presented a paper on tuning vLLM for coding agents.
- Background: They collected the most comprehensive Claude Code OS-level tracing dataset from their student labs.
- Core Finding: Existing serving frameworks like vLLM and SGLang cannot directly handle long-running agent workloads.
- Technical Detail: If KV cache is constrained, more CPU scheduling overhead is needed to predict cache reuse in long coding sessions.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- How Do You Catch Behavioral Regressions in LLM Agents Between Releases? — Beautiful_Belt_601 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- Run Firefox MCP on Android: Termux + ngrok tunnel tutorial — Nervous-Strain7544 · 2026-09-11