Open-source DKV framework cuts KV-cache memory for long-context local inference

Om_5000 · reddit · 2026-07-25

DKV (DifferentialKV) is an open-source framework for compressing KV cache in long-context local LLM inference.

The project focuses on reducing memory usage through anchor-based representations, joint low-rank compression, exact residual preservation, and sparse routed attention. It ships with a CLI, MLX backend, CUDA backend under validation, a technical report, and a fully open-source implementation. The author is looking for technical feedback and possible integrations with llama.cpp, vLLM, and SGLang.

Original post →

More from Infra

Infra channel →