Yandex Publishes 'The KV Cache as an Agent Runtime': Inference-Time Interactivity for LLMs
arpit_bhayani · x · 2026-09-07
Yandex Research published the full write-up behind the recent KV-cache buzz: 'The KV cache as an agent runtime.' Interactive models must take in new information mid-computation, revise trajectories, emit partial actions, and coordinate processes running at different rates — unavoidable for games, robots, live video, and OSes. While Wan-Streamer and Thinking Machines' Interaction Models solve this via retraining, new data formats, or streaming encoders, Yandex explores how much interactivity can be achieved purely at inference time by sharing and scheduling KV-cache state, letting pretrained LLMs observe, reason, and act concurrently without any training.
Related event: Yandex Turns KV Cache into an Agent Runtime for Real-Time LLM Interaction(3 posts)→
More from Infra
- Pod: NVIDIA's $12.93B all-cash Hugging Face acquisition and Anthropic's $35B compute bet — ryanshrout · 2026-09-08
- How much of daily life can run on an NVIDIA Jetson Nano? One blogger's open experiment — MaziyarPanahi · 2026-09-08
- Analyst flips bullish on DRAM: ~10% QoQ price hike seen in 4Q26, NAND softening — AccBalanced · 2026-09-08
- Compute financing risk will fall to hedge funds and commodity traders, not private credit — AccBalanced · 2026-09-08
- Solo dev ships Jenny, an MIT-licensed local LLM desktop app after 1.5 years — TangySword · 2026-09-08
- Cacheon launches GLM-5.3 kernel arena, paying up to 33 TAO daily to beat sglang — JosephJacks_ · 2026-09-08