Qwen3.8-Flash-Next on M3 Ultra: 559 t/s prompt processing, 31 t/s generation

rm-rf-rm · reddit · 2026-09-14

A Reddit user benchmarked Qwen3.8-Flash-Next (Q4KXL GGUF via unsloth) locally on an Apple M3 Ultra using llama.cpp through llama-swap with a 256K context. Measured with llama-benchy: 558.83 t/s prompt processing (pp1000) and 31.05 t/s token generation (tg500, peak 31.67), with 3.3s time-to-first-response. A useful data point for local deployment on Apple Silicon.

Original post →

More from Infra

Infra channel →