Ninfer Benchmark: 5090 Doubles Throughput for Qwen3 27B

Rollingsound514 · reddit · 2026-08-28

The author reports benchmark results running Qwen3 8 27B (nvfp4) on an RTX 5090 using the Ninfer engine. The setup reportedly more than doubles throughput compared to llama.cpp, peaking at 220 tokens/s and averaging in the 170s.

Specific Configuration:

The author praised the project's performance optimization.

Original post →

More from Infra

Infra channel →