vLLM vs ninfer for Qwen 3.8-27B on RTX 5090: a Reddit reality check
MaxKingCS · reddit · 2026-08-31
A Reddit user currently runs unsloth/Qwen3.8-27B-NVFP4 with 157k context via vLLM on an RTX 5090, and asks whether switching to the much-hyped ninfer is worth it. They also spotted a Hugging Face repo (gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090) claiming more context and better performance than the unsloth build — unverified and suspicious in their view — and want optimal configs for use with agent harnesses like Hermes agent.
More from Infra
- Local 8B Model Document Extraction Demo on iPhone 16 — Better_Comment_7749 · 2026-08-31
- Apple May Scrap 2027 Mobile HBM Plans Due to High Costs — power97992 · 2026-08-31
- Matmul Energy Efficiency Competition Challenges AlphaTensor Algorithm — yaroslavvb · 2026-08-31
- NVIDIA DGX Station offers data-center-class performance — SpendLucky1273 · 2026-08-31
- Comparing V100 and 5060Ti for Qwen 3.8 inference — ColorsOfCosmos · 2026-08-31
- Optimizing Qwen 3.8 27B with MTP and KV Cache Quantization — gabrielesilinic · 2026-08-31