Hong Kong team open-sources Agens Volundr 32B: only 18 of 72 layers keep a KV cache
ComfortableKindly507 · reddit · 2026-10-05
Blockway, a small Hong Kong team, released Agens Volundr 32B Preview (Apache-2.0), the first model on their own hybrid architecture designed to slash KV cache at long context:
- Architecture: 72 dense 32B layers — 54 KDA linear-attention layers (no KV cache), 17 BCSA compressed-sparse layers (exact 4K window; older context pooled 4:1, learned indexer reads top 512 blocks), 1 full-attention layer, plus Engram host-RAM n-gram memory and 4 residual streams (mHC). Only 18/72 layers keep a KV cache; 262K context.
- Speed: BF16 on two 48GB GPUs, decode barely degrades from 1K to 128K (25.1→23.9 tok/s); INT4 fits one 48GB card (31.7GiB). DFlash2 drafter: up to 3.6x on JSON/tool output, 2x code (not worth it above 8 concurrent users).
- Benchmarks: ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6); level on MMLU-Pro/GPQA; weak on agent tasks (tau2 74.2, SWE-bench subset 44) — the v1 focus.
- Limits: needs their sglang build; no GGUF/llama.cpp yet; still training.
Docker images and weights on GitHub/Hugging Face.
More from Models
- Speculation: OpenAI's next model could tackle long-context and KV cache costs — haider1 · 2026-10-05
- Japan AISI Evaluates Claude Opus 4.8 Cyber Skills: One pc_control Case, No Full T1 — HaydnBelfield · 2026-10-05
- Local Models Actually Beat Claude Opus 5.5 on Some Tasks in Hands-On Test — stefanjblos · 2026-10-05
- Dev reports Opus 5.5 'got dumber today,' speculating a new model update is imminent — TejasKumar_ · 2026-10-05
- Diffusion LMs Match Autoregressive Baselines on Math and Code After Pre-training — ricklamers · 2026-10-05
- Reuters study: near-identical LLM scorers overlap only 0.66-0.84 when candidates are reordered — thomsonreuters · 2026-10-05