Running a 27B Model on 16GB VRAM: NInfer 4080 Hits 262 tok/s Decode
roofkid · reddit · 2026-10-04
A developer with 20 years of software engineering experience released NInfer 4080, running the ISTA-DASLab-Qwen-3.8-27B-GSQ quant on an RTX 4080's 16GB VRAM at 100k context with 2720 tok/s peak prefill and 262 tok/s generation (GitHub: roofkid/ninfer-4080).
Key approaches
- DFlash2 speculative decoding + MTP3 to unlock performance general-purpose engines (llama.cpp/vllm) leave on the table — initial memory throughput of 200 GB/s vs a 720 GB/s theoretical max was jaw-dropping
- KV quantization trades accuracy for context; MBPP stays at 90-92%, HumanEval at 95-96%, with degradation within measurement noise
- Development delegated largely to DeepSeek V4.1 Flash ($13 of credits), Pi as harness, shipped as a Docker image for easy reproduction
At 98K context: 1895 tok/s prefill, 212 tok/s decode. The author notes general-purpose engines sacrifice far more performance for compatibility than expected.
More from Infra
- Why AI labs won't push on-device models: cloud inference is their business — AccBalanced · 2026-10-04
- uv fork runs dev sessions ~33% faster with 40-60% less disk, author still unsure it's enough — mitsuhiko · 2026-10-04
- KohakuFA: Blackwell flash attention kernel fixes 10x-1000x gradient bugs in FA4/cuDNN — bdsqlsz · 2026-10-04
- RWKV-7 G1k ships: pure-RNN reasoning with no KV cache, 16M-state 13B model — cephaloform · 2026-10-04
- NeoCloud Summit 2026 lands in SF Oct 8, gathering the GPU-native cloud ecosystem — AccBalanced · 2026-10-04
- Modal engineer built Gang Scheduler on K8s Reconciler model — it just worked — emilyzsh · 2026-10-04