Custom RTX 5090 inference stack claims 550–720 tokens/s on Qwen 3.6 35B
BringTea_666 · reddit · 2026-07-28
A custom RTX 5090 build claims 550–720 tokens/s on Qwen 3.6 35B
A Reddit user says Ninfer, a purpose-built inference engine for RTX 5090, can run Qwen 3.6 35B at roughly 550–720 tokens/s on Windows, with performance varying by task.
What the post says
- The user previously needed batched or parallel agent setups to reach similar throughput.
- With Ninfer, they get that speed in a single instance.
- The setup reportedly supports full 250k context.
- The project is described as a custom build for RTX 5090 and currently only supports Qwen 3.6 27B and 35B.
- The linked GitHub repository is Linux-first, though the poster says it can be made to work on Windows.
The headline claim is that the system feels close to Cerebras-like speeds on consumer hardware.
More from coding & agent
- Dev claims 20k more commits coming: Opus 5.5 and GPT-6 Sol supercharge his output — doodlestein · 2026-09-23
- A JEV-powered Wireshark classifier accidentally uncovered real backdoors on a home network — multiply_matrix · 2026-09-23
- 299 real intents tested: classifier routing trails GLM-4-Flash by 3 points but is 6.5x faster — Sufficient_Flower860 · 2026-09-23
- OpenExecutive: open-source virtual executive team of 8 specialist AI agents hits 5.1k GitHub stars — tom_doerr · 2026-09-23
- Framer launches Skills: teach your design agent reusable workflows, design systems and CMS rules — soleio · 2026-09-23
- Cursor, OpenAI and Anthropic shipped coordinator-agent fleets in one week, but the review bottleneck stays — omidfarhang · 2026-09-23