Qwen3.8-Flash-Next on M3 Ultra: 559 t/s prompt processing, 31 t/s generation
rm-rf-rm · reddit · 2026-09-14
A Reddit user benchmarked Qwen3.8-Flash-Next (Q4KXL GGUF via unsloth) locally on an Apple M3 Ultra using llama.cpp through llama-swap with a 256K context. Measured with llama-benchy: 558.83 t/s prompt processing (pp1000) and 31.05 t/s token generation (tg500, peak 31.67), with 3.3s time-to-first-response. A useful data point for local deployment on Apple Silicon.
More from Infra
- Rumble's Neocloud Pivot Lands $13.7B Compute Deal With Anthropic — andersonbcdefg · 2026-09-14
- Running an LLM agent on a 512MB board with decoupled memory and live cross-machine migration — D777Castle · 2026-09-14
- Maia 200 hits ~12 TFLOP/s FP4 in 1mm²: density should be a first-class goal — thoefler · 2026-09-14
- Dev's 24/7 self-hosted AI stack: OpenWebUI, pidot, Tailscale, GLM and DeepSeek — andfanilo · 2026-09-14
- Dream Photonics' laser integration render caught mirroring the whole chip image — jwt0625 · 2026-09-14
- Engineer questions the inline PTX hype DeepSeek sparked: hand-written PTX isn't a proxy for performance — mike64_t · 2026-09-14