Ninfer hits 130-180 tok/s on Blackwell, dramatically faster than llama-server

swagonflyyyy · reddit · 2026-10-02

A Reddit user benchmarked the new Ninfer inference backend on a Blackwell MaxQ (qwen3.8-27b-nvfp4, Dflash2, MTP-7), reaching 130-180 tok/s with barely any slowdown in parallel decoding and concurrency on par with vLLM. They migrated their daily driver from llama-server on Windows 10 and praised its fit for agentic vibecoding, noting missing advanced sampling params and model support for now.

Original post →

More from Infra

Infra channel →