Yandex researchers propose KV cache as an agent runtime, demo Qwen3.8 playing DOOM interactively
_puhsu · reddit · 2026-09-07
Yandex's research team published a blog post proposing to modify the model's inference state (KV cache) directly as an agent runtime for better interactivity and responsiveness. The idea builds on their prior papers Hogwild! Inference and AsyncReasoning, and the post previews upcoming work: a Qwen3.8-27B agent playing a DOOM environment interactively with similar techniques.
Their core argument: inference/runtime design may be an under-explored axis of agent capability — existing harnesses are too abstract and swapping models is too costly, so something in between may be needed.
More from coding & agent
- Google AI Search MCP server adds grounded web search via Vertex AI or Gemini — modelcontextprotocol · 2026-09-07
- Asterwise launches Vedic astrology MCP server with real ephemeris calculations — modelcontextprotocol · 2026-09-07
- CubeSandbox v0.7.0 ships cross-node pause & resume for AI agent sandboxes — HeyAmit_ · 2026-09-07
- GPT-6 Astra turns one prompt into a 1440p Colosseum 3D video via Blender MCP — Scobleizer · 2026-09-07
- awesome-ai-apps: open-source repo with 132 LLM app examples, agents and RAG demos — Arindam_1729 · 2026-09-07
- Why 88-95% of enterprise AI agent pilots never ship — and what working teams do differently — ankitsharma112 · 2026-09-07