Qwen 350K Context Tested on M5 Max: Performance and Quality

Artistic_Okra7288 · reddit · 2026-08-30

Benchmark of 2-bit quantized Qwen3.8-Flash-Next running via llama.cpp on a 128GB M5 Max, utilizing a 358K-token context slot. Results show that while prefix reuse keeps prefills fast, a cold prefill after idle took 333s. Decode throughput tapers from 30-35 t/s at small context to 11.5 t/s at 169K tokens. The model performed strongly up to 100K tokens but exhibited role confusion at extreme depths, likely due to 2-bit quantization or preview model limitations.

Original post →

More from Infra

Infra channel →