Open-source DKV framework cuts KV-cache memory for long-context local inference
Om_5000 · reddit · 2026-07-25
DKV (DifferentialKV) is an open-source framework for compressing KV cache in long-context local LLM inference.
The project focuses on reducing memory usage through anchor-based representations, joint low-rank compression, exact residual preservation, and sparse routed attention. It ships with a CLI, MLX backend, CUDA backend under validation, a technical report, and a fully open-source implementation. The author is looking for technical feedback and possible integrations with llama.cpp, vLLM, and SGLang.
More from Infra
- Open-source DKV cuts KV-cache memory for local long-context LLM inference — Om_5000 · 2026-07-25
- Intel consumer motherboards can break PCIe P2P on multi-GPU AI rigs — Arli_AI · 2026-07-25
- Scale-out networking doesn’t need CPO, slide argues; compute still dominates power use — zephyr_z9 · 2026-07-25
- Kimi weights could turn the debate into hardware economics versus V4 — teortaxesTex · 2026-07-25
- Jensen Huang hands Elon Musk a desk-size DGX Spark at Starbase — XFreeze · 2026-07-25
- Xpeng starts pilot production of humanoid robots as Anthropic eyes in-house chips — 创业邦 · 2026-07-25