Blockway Ships Agens Volundr 32B: Only 18 Layers Keep KV Cache for Long Context
lmoroney · x · 2026-10-08
Hong Kong's Blockway released Agens Volundr 32B Preview on Hugging Face under Apache 2.0, designed to slash KV-cache memory for long conversations. Of 72 layers, only 18 keep a KV cache: 54 use Kimi Delta Attention (fixed-size linear state), 17 use a 4,096-token sliding window plus compressed far field, and 1 uses full attention. On two 48GB GPUs, decode speed drops only from 25.1 to 23.9 tok/s as context grows from 1K to 128K; an INT4 build fits on one 48GB card. Caveats: it hit repetition loops in 36 of 50 SWE-bench tasks and currently requires Blockway's own sglang build (no vLLM plugin yet).
More from Infra
- WSJ: Broadcom arranging $50B+ financing for OpenAI's custom chips under secret Nexus program — rohanpaul_ai · 2026-10-08
- Reddit asks: have AI scaling laws hit their limit, or is the compute buildout just starting? — StupidDialUp · 2026-10-08
- TRIAGE stabilizes native NVFP4 RL training, hits full-precision quality at 2.3x throughput — InfiX-ai · 2026-10-08
- WSL Containers Now Generally Available: Run Linux Containers Natively on Windows — pavandavuluri · 2026-10-08
- ai& Says It's Japan's Largest Dedicated Inference Provider, Teases Post-Training Offerings — DavidBennett__ · 2026-10-08
- Microsoft unveils Surface Laptop Ultra with Nvidia RTX Spark SoC from $2,599 — Ars Technica AI · 2026-10-08