1M Token Context on Single RTX 3090 Achieved via KVarN Quantization
Anbeeld · reddit · 2026-08-10
Developer Anbeeld shared an extreme local LLM deployment test: a user successfully loaded nearly 1 million tokens of context for a 17GB Qwen 3.5 35B A3B model on a single RTX 3090 (24GB VRAM).
The test wasn't just about the server staying alive; the user successfully extracted 7 "needles" (precise information) placed in various parts of the massive text, proving effective long-context understanding under extreme quantization.
This breakthrough utilizes the BeeLlama.cpp fork and Huawei's KVarN (Variance-Normalized KV-Cache Quantization). Compared to standard 4-bit quants, KVarN demonstrates superior precision at ultra-low bitrates, fundamentally changing expectations for low-bit KV cache quantization performance.
More from Infra
- Intel Announces $15B Stock Offering to Capitalize on AI Compute Demand — ryanshrout · 2026-08-10
- Hybrid Cloud + Local Model Architecture Shows Promise; Muse Spark 1.2 Open-Source Release Imminent — jack_w_rae · 2026-08-10
- 30B Model Hits 114 tps with Tuned Quants, Targeting <16G VRAM — tokenbender · 2026-08-10
- Beyond Compute: AI Data Centers Face the Power Scarcity Bottleneck — ingliguori · 2026-08-10
- Discovered Materials Raises $9M to Hunt for Novel Chip Cooling Materials — TechCrunch AI · 2026-08-10
- Choosing MiniMax H3 Quantization for RTX 5090: int8 vs nvfp4 — Zerozone000 · 2026-08-10