Qwen 350K Context Tested on M5 Max: Performance and Quality
Artistic_Okra7288 · reddit · 2026-08-30
Benchmark of 2-bit quantized Qwen3.8-Flash-Next running via llama.cpp on a 128GB M5 Max, utilizing a 358K-token context slot. Results show that while prefix reuse keeps prefills fast, a cold prefill after idle took 333s. Decode throughput tapers from 30-35 t/s at small context to 11.5 t/s at 169K tokens. The model performed strongly up to 100K tokens but exhibited role confusion at extreme depths, likely due to 2-bit quantization or preview model limitations.
More from Infra
- Azure Linux 4.0 Desktop Concept: PowerShell, Edge, and Copilot Pre-installed — unixterminal · 2026-08-30
- Jensen Huang: Built GPU tech first, found endless problems from graphics to molecular dynamics — r0ck3t23 · 2026-08-30
- How to build an LLM inference engine from scratch: 5-layer architecture — glenbeer · 2026-08-30
- Huaqin expects super node revenue to exceed 10B RMB in 2H 2026 — zephyr_z9 · 2026-08-30
- Nvidia is generating $1 billion a day, a business scale deemed absurd years ago — shauntrennery · 2026-08-30
- Krishnan: space data centers streaming intelligence into our nerve centers will feel shockingly normal — sebkrier · 2026-08-30