Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now?

jankeydankey · reddit · 2026-09-23

A user shared his local baseline: Qwen 3.6 35B-A3B at Q6 runs at a steady 50 tokens/s on a 128GB Strix Halo APU under Nobara with llama.cpp — stable enough that he stopped tinkering for months. Returning to the scene, he asks whether any newer model offers more capability at the same speed (or vice versa), and what the current best setup for 128GB Strix Halo is.

Original post →

More from Infra

Infra channel →