Qwen 3.8 Flash Next hits 3.1k tok/s prefill on M5 Ultra

GabGarrett · x · 2026-09-22

mweinbach had astra optimize Qwen 3.8 Flash Next on an M5 Ultra, now reaching 3.1k tok/s prefill. GabGarrett notes that while Spark fans rushed to dunk on the Ultra, the machine may turn out to be a beast — a promising sign for Apple silicon unified memory in local LLM inference.

Original post →

More from Infra

Infra channel →