Qwen 3.8 27B hits 96 t/s decode with 110k context on a single 16GB RTX 5080 via NInfer

Kernoriordan · reddit · 2026-10-07

A developer improved on their earlier 75 t/s llama.cpp setup by running Qwen 3.8 27B with NInfer v1.5 on a 16GB RTX 5080 (Ubuntu 24.04 via WSL2), achieving 90–110 t/s decode with 110,592 tokens of context.

Real-world numbers (32 Zoo Code coding requests)

Key config

Gotchas

No controlled quality comparison vs GGUF yet.

Original post →

More from Infra

Infra channel →