User asks what hardware can push Qwen3.6 35B to 1,000+ prefill throughput
Mrinohk · reddit · 2026-07-21
A Reddit user asks what hardware people are using to hit 1,000+ prefill and 100+ decode throughput on Qwen3.6 35B at Q4.
- Their current setup is an RX 6600 XT, Ryzen 7 5700X, and 32 GB DDR4.
- They report about 270–300 tokens/s on prefill and roughly 30 tokens/s on decode using llama.cpp with ROCm on CachyOS.
- They say tuning GPU layers or placing experts in system memory usually made performance worse on their machine.
- The poster argues public hardware benchmarks are unreliable and wants real-world configs from people reaching much higher speeds.
More from Infra
- SkyPilot exits stealth with $20M seed round and an AI compute platform for fragmented clouds — jfiance · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22