NInfer vs llama.cpp vs vLLM: rigorous quality and speed test of Qwen3.8-27B NVFP4 on one RTX 5090
bengizmoed · reddit · 2026-09-05
A developer shares a rigorously controlled comparison of three inference engines serving Qwen3.8-27B NVFP4 on a single RTX 5090 (32GB, eGPU), evaluated on 6 tiers of real production workloads — relevance classification, needle retrieval at 64K-240K context, multi-transcript QA, reasoning, structured extraction, and agent tool replay — with paired prompts, seeds, gold labels, and cluster-bootstrap CIs. A 4-model review panel (Codex/GPT-5, DeepSeek V4 Pro, Grok 4.5, Kimi K3) audited the methodology and caught 10 issues including 4 mislabeled gold items.
Key findings:
- Quality is statistically indistinguishable across all three engines; needle retrieval is 100% on every item that fits
- NInfer's retrieval edge comes entirely from context capacity (240K vs llama.cpp's 196K), not quality
- All engines fail tool replay equally; llama.cpp is limited to parallel=1
- NInfer offers MTP3 (76% acceptance) with x2 concurrency; VRAM use is 29.6-31.6GB across engines
Takeaway: pick based on concurrency and context needs, not quality fears.
More from Infra
- Who Needs a GB300? Researcher Makes a Movie on a $500 16GB GPU — francoisfleuret · 2026-09-05
- Minima quantizes all 496 layers of Qwen3.8-27B to NVFP4 W4A4, matching BF16 at 2.9x smaller — pbaylies · 2026-09-05
- 19 latency patterns to cut non-model latency in AI applications — bibryam · 2026-09-05
- Astra reportedly trained on 100,000 GPUs, a staggering compute scale — sudoraohacker · 2026-09-05
- GPT-6 reportedly trained on just 100K B200/300 GPUs at a single Texas facility — round · 2026-09-05
- gfx906-llama-cpp Fork Boosts MI50/MI60 Inference +23% Prefill, Fits 250k Context on 40GB — milpster · 2026-09-05