Single RTX 4090 runs Qwen 27B with up to 350K-token context via KV cache quantization
UDPSendToFailed · reddit · 2026-08-17
Developer UDPSendToFailed updated their fork of the NInfer inference engine, adding an rk2v4-e8 quantization option for the KV cache. Running a Qwen 27B model on a single RTX 4090, it can handle a context window of 250K-350K tokens entirely in VRAM without spilling into system RAM; the actual ceiling depends on configuration options like vision and MTP.
For lower-context runs, the author also optimized generation speed to roughly 80-160 tokens/s on repetitive workloads such as code and math. The project is open-sourced on GitHub (UDPSendToFailed/ninfer-4090), and the author welcomes reports of runtime issues.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24