Dev hacks llama.cpp for NVFP4 KV cache, runs Qwen3 27B at 262k context across two GPUs

comperr · reddit · 2026-09-19

A Reddit user got NVFP4 KV cache working in a patched llama.cpp (heterogeneous port of an ollama branch), running a self-made Q4KM quant of Qwen3 27B at 262k context split across an RTX 5090 and RTX 3090 Ti. Key technical points:

Directly useful for anyone doing long-context local deployment or quantization work on consumer multi-GPU setups.

Original post →

More from Infra

Infra channel →