Benchmark: NInfer Boosts Qwen3.6 Prefill Speed Over 2x vs llama.cpp

tat_tvam_asshole · reddit · 2026-08-01

A developer benchmarked NInfer (NVFP4) against llama.cpp (Q4KXL) running the Qwen3.6-27B model on a power-limited RTX Pro 6000.

NInfer is highly optimized for sm120 (targeting future 5090 GPUs), and this test provides valuable community benchmarks for current hardware.

Original post →

More from Infra

Infra channel →