Qwen 3.8 Flash Next hits 3.1k tok/s prefill on M5 Ultra
GabGarrett · x · 2026-09-22
mweinbach had astra optimize Qwen 3.8 Flash Next on an M5 Ultra, now reaching 3.1k tok/s prefill. GabGarrett notes that while Spark fans rushed to dunk on the Ultra, the machine may turn out to be a beast — a promising sign for Apple silicon unified memory in local LLM inference.
More from Infra
- Raspberry Pi locks devices to original RAM size, blocking aftermarket memory upgrades — ngxson · 2026-09-22
- fal's H3 Max generates 5 seconds of frontier-quality video in just 3 seconds — gorkem · 2026-09-22
- NVIDIA's EPD Disaggregation Cuts Multimodal TTFT Up to 5x, E2E Latency 7x — dl_weekly · 2026-09-22
- Running MiniMax H3 locally on a 16GB Mac: 8-10s clips in 15-20 minutes — coberholzer · 2026-09-22
- A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4 — willcb · 2026-09-22
- AMD's market cap went from $2B to $1T in the 11 years since Lisa Su became CEO — xiaosun86 · 2026-09-22