Ninfer Benchmark: 5090 Doubles Throughput for Qwen3 27B
Rollingsound514 · reddit · 2026-08-28
The author reports benchmark results running Qwen3 8 27B (nvfp4) on an RTX 5090 using the Ninfer engine. The setup reportedly more than doubles throughput compared to llama.cpp, peaking at 220 tokens/s and averaging in the 170s.
Specific Configuration:
- Model: qwen3827bnvfp4.ninfer
- Context/KV: 240k capacity, fp8 KV dtype, 16GiB host KV.
- Speculative Decoding: Enabled mtp and lm-head-draft with 3 draft tokens.
- Vision: Enabled vision and media-live (2GiB).
The author praised the project's performance optimization.
More from Infra
- Google, NVIDIA Back Open Source Project to Optimize LLM Inference on Kubernetes — SumitGup · 2026-08-28
- Chinese AI Models Scale Up: GLM 5.3 Flash Trains on 30T Tokens — teortaxesTex · 2026-08-28
- Australia Minister: No Fossil Fuel Carve-out for Datacenters — nordicinst · 2026-08-28
- ZED Camera priced at $500? DIY alternative costs just $150 — _William_F_ · 2026-08-28
- KOTOR Remaster Path Tracer Integrates DLSS 4.5 RR — Michael_Moroz_ · 2026-08-28
- Tutorial: Train a Raspberry Pi to Read Gas Meter Automatically with Neural Network — JeremyCMorgan · 2026-08-28