Custom RTX 5090 inference stack claims 550–720 tokens/s on Qwen 3.6 35B
BringTea_666 · reddit · 2026-07-28
A custom RTX 5090 build claims 550–720 tokens/s on Qwen 3.6 35B
A Reddit user says Ninfer, a purpose-built inference engine for RTX 5090, can run Qwen 3.6 35B at roughly 550–720 tokens/s on Windows, with performance varying by task.
What the post says
- The user previously needed batched or parallel agent setups to reach similar throughput.
- With Ninfer, they get that speed in a single instance.
- The setup reportedly supports full 250k context.
- The project is described as a custom build for RTX 5090 and currently only supports Qwen 3.6 27B and 35B.
- The linked GitHub repository is Linux-first, though the poster says it can be made to work on Windows.
The headline claim is that the system feels close to Cerebras-like speeds on consumer hardware.
More from coding & agent
- MyClaw adds $100 to $200 in AI credits on annual Pro, Max and Ultra plans — manishkhosiya · 2026-07-28
- MyClaw lets subscribers try OpenClaw and Hermes Max plans for 30 days — manishkhosiya · 2026-07-28
- MyClaw adds connectors for Gmail, Slack, GitHub, Shopify and more — manishkhosiya · 2026-07-28
- MyClaw pitches 30-second setup and 24/7 hosting for AI agents — manishkhosiya · 2026-07-28
- Microsoft open-sources a governance toolkit for autonomous AI agents — microsoft · 2026-07-28
- A Python repo turns technical book PDFs into Claude Code skills — virgiliojr94 · 2026-07-28