Running a 27B Model on 16GB VRAM: NInfer 4080 Hits 262 tok/s Decode

roofkid · reddit · 2026-10-04

A developer with 20 years of software engineering experience released NInfer 4080, running the ISTA-DASLab-Qwen-3.8-27B-GSQ quant on an RTX 4080's 16GB VRAM at 100k context with 2720 tok/s peak prefill and 262 tok/s generation (GitHub: roofkid/ninfer-4080).

Key approaches

At 98K context: 1895 tok/s prefill, 212 tok/s decode. The author notes general-purpose engines sacrifice far more performance for compatibility than expected.

Original post →

More from Infra

Infra channel →