Qwen 3.8 27B vs Qwen Flash Next on an M3 Max: near-identical feel, faster prefill on 27B
Zeeplankton · reddit · 2026-09-07
The poster runs Qwen 3.8 27B and Qwen Flash Next locally on an M3 Max 96GB: both feel largely identical, though 27B prefills faster; they ask whether anyone is improving prefill performance in MLX.
Side question: is anyone building a harness that works with reasoning fully off, inspired by JetBrains revealing Junie runs Qwen 3.6 with reasoning disabled entirely.
More from Infra
- Compute financing risk will fall to hedge funds and commodity traders, not private credit — AccBalanced · 2026-09-08
- Solo dev ships Jenny, an MIT-licensed local LLM desktop app after 1.5 years — TangySword · 2026-09-08
- Cacheon launches GLM-5.3 kernel arena, paying up to 33 TAO daily to beat sglang — JosephJacks_ · 2026-09-08
- Johns Hopkins report urges strategic power islanding to cut blackout risk — philvenables · 2026-09-08
- HBF math: matching H200 bandwidth needs ~4,900 concurrent NAND planes — lauriewired · 2026-09-08
- BMO: record US power demand isn't lifting gas use — renewables and batteries are filling the gap — aronchick · 2026-09-08