Microsoft Analyzes 13.5M Copilot Sessions: Why Agent Scheduling Differs from Chat
rohanpaul_ai · x · 2026-08-09
A new Microsoft paper analyzing 13.5 million GitHub Copilot production sessions reveals that infrastructure scheduling for coding agents must fundamentally differ from standard chat requests.
- Call Distribution: 87% of LLM calls are initiated autonomously by the agent rather than the user. A single prompt fans out into chains of model calls, tool actions, and retries.
- KV Cache Behavior: Cache hit rates depend heavily on workflow position. Within a single turn, hit rates rise from 45% on the first call to 92–94% by the third. However, it plummets to 55% at turn boundaries and 8% when switching models.
- Idle Times: Median KV-cache idle time is 1.2s within a turn versus 172s across turns. Container idle time jumps from 5.8s to 243s.
- Scheduling Implications: Using turn- and session-level features, the paper's lightweight predictor captures 86–90% of total idle time. Infrastructure should schedule based on workflow state to optimize cache offloading and container reclamation.
More from Infra
- JPMorgan: AI Chips May Need Over $2T in 5 Years, Fueling Bond Markets — rohanpaul_ai · 2026-08-09
- Stanford's Yi Cui Outlines Three Battery Tech Paths for AI Data Centers — FinanceYF5 · 2026-08-09
- DeepSeek V4 Flash Local Quantization Benchmark on SlopCodeBench — corruptbytes · 2026-08-09
- Beef vs. Data Centers: The Double Standard in AI Water Consumption Debates — KrustyKrabFormula_ · 2026-08-09
- Replace ChatGPT Plus with Local Models: A 5-Step Guide — Aiden_Tech_Ai · 2026-08-09
- Nvidia's Rubin Ultra Shifts from HBM to Optical Interconnects, Altering Market Dynamics — zephyr_z9 · 2026-08-09