456 tok/s Qwen 3.8 on Modded RTX 2080 Ti via NInfer Port

xrailgun · reddit · 2026-08-21

The author ported and tuned the NInfer C++/CUDA inference engine, originally designed for RTX 5090/Blackwell, to NVIDIA Turing architecture (sm75), specifically targeting a modded 22GB RTX 2080 Ti.

Benchmarks (Qwen 3.8-27B MTP3, W8A16, Q8 KV):

The project includes OpenAI/Anthropic-compatible HTTP serving (streaming, function calling) and is open-sourced on GitHub.

Original post →

More from Infra

Infra channel →