1M Token Context on Single RTX 3090 Achieved via KVarN Quantization

Anbeeld · reddit · 2026-08-10

Developer Anbeeld shared an extreme local LLM deployment test: a user successfully loaded nearly 1 million tokens of context for a 17GB Qwen 3.5 35B A3B model on a single RTX 3090 (24GB VRAM).

The test wasn't just about the server staying alive; the user successfully extracted 7 "needles" (precise information) placed in various parts of the massive text, proving effective long-context understanding under extreme quantization.

This breakthrough utilizes the BeeLlama.cpp fork and Huawei's KVarN (Variance-Normalized KV-Cache Quantization). Compared to standard 4-bit quants, KVarN demonstrates superior precision at ultra-low bitrates, fundamentally changing expectations for low-bit KV cache quantization performance.

Original post →

More from Infra

Infra channel →