Qwen3.8-Flash-Next Benchmarked Locally: First Model to Break 94% on 128GB Mac
tolitius · reddit · 2026-08-27
A developer tested Qwen3.8-Flash-Next on an M4 Max 128GB Mac Studio with oMLX and llama.cpp — the first model this year to break 94% on their custom cupel benchmark.
- The qwen4exp architecture isn't supported yet, so oMLX K/V caching had to be disabled; the full 4-bit quant takes 100GB, a tight fit
- On a mixed coding/general-knowledge/science benchmark, Qwen 3.8 27B led in coding but lost on general knowledge to Gemma 31B and Qwen 3.6
- Best quant: pipenetwork's MLX mixed-4/8bit build (perplexity 4.5286 vs 4.4708 bfloat16); Unsloth's GGUF UD-IQ4XS is weaker but smaller
The author plans to add more coding toolchains (pi/opencode) as models get too good to differentiate.
More from Models
- Perplexity Computer Launches: Local/Cloud Support, Claims 85.4% Benchmark — ChrisUniverse · 2026-08-27
- Critique of Claude style: Verbose but information-dense — TheZvi · 2026-08-27
- GLM-5.3 Flash review: Fixes 13 bugs with high cost-performance — PawelHuryn · 2026-08-27
- Opus 5 writing regression confirmed by benchmarks — gleech · 2026-08-27
- Gemini 3.7 Flash beats benchmarks at 96% lower cost — DynamicWebPaige · 2026-08-27
- User Touts DeepSeek Reasoning Model Performance on Complex Logic Test — vista8 · 2026-08-27