vLLM vs ninfer for Qwen 3.8-27B on RTX 5090: a Reddit reality check

MaxKingCS · reddit · 2026-08-31

A Reddit user currently runs unsloth/Qwen3.8-27B-NVFP4 with 157k context via vLLM on an RTX 5090, and asks whether switching to the much-hyped ninfer is worth it. They also spotted a Hugging Face repo (gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090) claiming more context and better performance than the unsloth build — unverified and suspicious in their view — and want optimal configs for use with agent harnesses like Hermes agent.

Original post →

More from Infra

Infra channel →