456 tok/s Qwen 3.8 on Modded RTX 2080 Ti via NInfer Port
xrailgun · reddit · 2026-08-21
The author ported and tuned the NInfer C++/CUDA inference engine, originally designed for RTX 5090/Blackwell, to NVIDIA Turing architecture (sm75), specifically targeting a modded 22GB RTX 2080 Ti.
Benchmarks (Qwen 3.8-27B MTP3, W8A16, Q8 KV):
- Standard Autoregressive (MTP0): 25 tok/s.
- Speculative Decoding (MTP3, draft window=3): 456 tok/s (65% acceptance rate).
- VRAM Usage: 17.5 GiB total, leaving 4.5–5.0 GiB free for KV cache, enabling 128k context.
The project includes OpenAI/Anthropic-compatible HTTP serving (streaming, function calling) and is open-sourced on GitHub.
More from Infra
- Kubernetes CPU Limits Make Apps Slow and Costly: Proof and Experiments — JeremyCMorgan · 2026-08-21
- Productionizing AI Apps: OpenTelemetry, On-Call Agents, and Full Observability Workflow — Al_Grigor · 2026-08-21
- LLMRouter 2.0: Unified Infrastructure for LLM Routing Dev and Eval — youjiaxuan · 2026-08-21
- Run MiniMax H3 locally on 12GB GPUs: 15-second multi-shot ComfyUI template — vortis23 · 2026-08-21
- llama.cpp adds tensor split for LFM2/MoE, boosting inference performance significantly — pmttyji · 2026-08-21
- Used RTX 3090 purchase review: Local AI performance crushes 3060, Qwen 35B 20x faster — Yanzihko · 2026-08-21