Custom RTX 5090 inference stack claims 550–720 tokens/s on Qwen 3.6 35B

BringTea_666 · reddit · 2026-07-28

A custom RTX 5090 build claims 550–720 tokens/s on Qwen 3.6 35B

A Reddit user says Ninfer, a purpose-built inference engine for RTX 5090, can run Qwen 3.6 35B at roughly 550–720 tokens/s on Windows, with performance varying by task.

What the post says

The headline claim is that the system feels close to Cerebras-like speeds on consumer hardware.

Original post →

More from coding & agent

coding & agent channel →