600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000
CharlesStross · reddit · 2026-09-18
A Reddit user reports 600 tok/s on a single request running Qwen3.6 35BA3B with Ninfer on an RTX Pro 6000. Even if it uses 20x more tokens, it's still faster than many local models for read-and-find or brute-force coding tasks — 'not quite Cerebras' but a fun high-throughput local setup.
More from Infra
- Mystery trader drops $100M in premium on 2-week AI stock calls expiring Oct 2 — toptickcrypto · 2026-09-18
- First-ever PyTorch Day Japan lands in Tokyo on December 10, CFP open till Sept 27 — PyTorch · 2026-09-18
- Self-built Blackwell Colab pipeline generates a 2-hour MiniMax H3 movie for $6.96 — Interesting-Town-433 · 2026-09-18
- LaurieWired's CppCon keynote covers new memory hierarchies and how to prepare — lauriewired · 2026-09-18
- NVIDIA unpacks how CUDA's full stack powers specialized AI across finance, health and manufacturing — NVIDIA Developer · 2026-09-18
- Cadence sees India's EDA market doubling to $7.82B by 2031 — bookwormengr · 2026-09-18