Benchmark: NInfer Boosts Qwen3.6 Prefill Speed Over 2x vs llama.cpp
tat_tvam_asshole · reddit · 2026-08-01
A developer benchmarked NInfer (NVFP4) against llama.cpp (Q4KXL) running the Qwen3.6-27B model on a power-limited RTX Pro 6000.
- Prefill Throughput: NInfer is 2.17x to 3.53x faster across 8K to 256K context lengths.
- Generation Speed: NInfer is 1.14x faster for code generation and 1.28x faster for structured JSONL workloads.
- Memory: NInfer used about 2.9GB (10%) less GPU memory during full-context runs.
NInfer is highly optimized for sm120 (targeting future 5090 GPUs), and this test provides valuable community benchmarks for current hardware.
More from Infra
- What Hardware Do You Need to Run Full DeepSeek Locally? Community Weighs In — whoami-233 · 2026-08-01
- Lost Tacit Knowledge Threatens Rapid Nuclear Buildout for AI — gabriel1 · 2026-08-01
- Qualcomm Acquires Modular to Tackle AI Software Bottlenecks — clattner_llvm · 2026-08-01
- Elon Musk Declares >99.99% of AI Computing Will Eventually Move to Space — Polymarket · 2026-08-01
- CoreWeave Launches Sandboxes: An Execution Layer for Agentic AI — wandb · 2026-08-01
- Engine-Agnostic Rust LLM Gateway SMG Graduates from LightSeek Foundation — vllm_project · 2026-08-01