600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000

CharlesStross · reddit · 2026-09-18

A Reddit user reports 600 tok/s on a single request running Qwen3.6 35BA3B with Ninfer on an RTX Pro 6000. Even if it uses 20x more tokens, it's still faster than many local models for read-and-find or brute-force coding tasks — 'not quite Cerebras' but a fun high-throughput local setup.

Original post →

More from Infra

Infra channel →