Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now?
jankeydankey · reddit · 2026-09-23
A user shared his local baseline: Qwen 3.6 35B-A3B at Q6 runs at a steady 50 tokens/s on a 128GB Strix Halo APU under Nobara with llama.cpp — stable enough that he stopped tinkering for months. Returning to the scene, he asks whether any newer model offers more capability at the same speed (or vice versa), and what the current best setup for 128GB Strix Halo is.
More from Infra
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23