gufo-Qwen3.6-35B hits 3095 tok/s prefill, 190 tok/s decode on Strix Halo
nubela · reddit · 2026-10-03
A Reddit user shared local inference benchmarks for gufo-Qwen3.6-35B-A3B-Q6dense on AMD's Strix Halo: 3095 tok/s prefill and 190 tok/s decode. The model is open-sourced on GitHub, a useful datapoint for consumer-grade local LLM deployment.
More from Infra
- Qwen 177B at 11-15 tok/s on a Single RTX 5070 12GB: Expert Streaming Deep Dive — ayobluestarr · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- SemiAnalysis: Nvidia's custom NVHBM frees ~25% more compute die area on Feynman — zephyr_z9 · 2026-10-03
- llama.cpp PR Halves Indexer Score Memory for Qwen Flash, Cutting VRAM Use — jacek2023 · 2026-10-03
- Garage server farms return: GPU and power shortages reverse the AWS era in Palo Alto — bookwormengr · 2026-10-03
- AI maxi calls GPU price hike a bubble peak, plans to buy cheap cards after burst — AIFlow_ML · 2026-10-03