Ninfer hits 130-180 tok/s on Blackwell, dramatically faster than llama-server
swagonflyyyy · reddit · 2026-10-02
A Reddit user benchmarked the new Ninfer inference backend on a Blackwell MaxQ (qwen3.8-27b-nvfp4, Dflash2, MTP-7), reaching 130-180 tok/s with barely any slowdown in parallel decoding and concurrency on par with vLLM. They migrated their daily driver from llama-server on Windows 10 and praised its fit for agentic vibecoding, noting missing advanced sampling params and model support for now.
More from Infra
- US data center construction hits record $85B/yr, up 513% since ChatGPT launch — HaydnBelfield · 2026-10-02
- HBM shortage to persist for years, analyst pushes back on Wall Street's 2029 memory crash bet — JOBhakdi · 2026-10-02
- Hugging Face adds Train Models docs: fine-tune on one A10G for ~$0.10 in minutes — vanstriendaniel · 2026-10-02
- Gemini 4 may roll out to Ultra users next week as Google's TPU capacity reportedly runs tight — haider1 · 2026-10-02
- Banks tighten data center lending as Beff bets thermodynamic computing can fix AI power crunch — beffjezos · 2026-10-02
- Recovering 4.1 GiB of hidden RAM on NVIDIA DGX Spark, 3.6x bigger KV pool — marian_nmt · 2026-10-02