2.3x faster Qwen3.8 27B on RTX 5090: ninfer vs llama.cpp benchmarked across 4 setups
theexile1337 · reddit · 2026-10-03
- The author benchmarked ninfer vs llama.cpp in 4 configurations running Qwen3.8 27B on an RTX 5090 with 120k context and xhigh thinking, then stress-tested each generated billiards game with 2,400 random shots plus scripted rule scenarios.
- Speed: non-NVFP4 ninfer with MTP hit 147 t/s (7m40s), while llama.cpp jumped from 68 t/s to 141 t/s once MTP was enabled—suggesting the "ninfer is fast" story is largely "MTP is fast." NVFP4 had the highest t/s (155) but the most thinking tokens, finishing only third in wall-clock time.
- Quality: all four games ran, but only the ninfer NVFP4 build could be legally won; the others had win-condition logic bugs (one an inverted one-line check).
- Caveats: one run per setup; NVFP4 and Q4KM quantizations aren't directly comparable. The author picked the non-NVFP4 ninfer build as most polished.
More from Infra
- DGX Spark shortage derails $8,800 donation plan as buyer can't find stock — cyrus_zei · 2026-10-03
- Perplexity to vertically integrate agentic infra on NVIDIA's Vera CPU, ditches x86 — AravSrinivas · 2026-10-03
- GPT-6 Astra burns 30 minutes of compute on oversized tasks, then dumps all progress at the limit — CautiousMagazine3591 · 2026-10-03
- AirLLM runs a 70B LLM on a single 4GB GPU by streaming layer weights from disk — techNmak · 2026-10-03
- SemiAnalysis asks if GPUs are money printers: TPU v7, Vera Rubin vs Blackwell — AccBalanced · 2026-10-03
- gufo Ships Pre-packaged Windows Build for AMD Strix Halo, 40 tps on Agentic Workloads — hiImMate · 2026-10-03