NInfer vs llama.cpp vs vLLM: rigorous quality and speed test of Qwen3.8-27B NVFP4 on one RTX 5090

bengizmoed · reddit · 2026-09-05

A developer shares a rigorously controlled comparison of three inference engines serving Qwen3.8-27B NVFP4 on a single RTX 5090 (32GB, eGPU), evaluated on 6 tiers of real production workloads — relevance classification, needle retrieval at 64K-240K context, multi-transcript QA, reasoning, structured extraction, and agent tool replay — with paired prompts, seeds, gold labels, and cluster-bootstrap CIs. A 4-model review panel (Codex/GPT-5, DeepSeek V4 Pro, Grok 4.5, Kimi K3) audited the methodology and caught 10 issues including 4 mislabeled gold items.

Key findings:

Takeaway: pick based on concurrency and context needs, not quality fears.

Original post →

More from Infra

Infra channel →