Qwen3.8 pMLX Engine: Runs on 12GB RAM at 10 tok/s
EyalToledano · x · 2026-09-01
Qwen3.8-Flash-Next is launching with a custom pMLX engine supporting tiered models and dynamic quantization (bf16/q8/q4/q3). Features include on-the-fly expert pruning, adjustable n-gram streaming, and NVME offloading. Benchmarks show code editing at 58 tok/s sustained on 32k context (Q4), and a low-memory config running at 10 tok/s on just 12GB RAM (Q3).
More from Infra
- RTX 6000 Blackwell crashes under load, points to firmware bug — AIFlow_ML · 2026-09-01
- Adani: AI competitive advantage shifts to clean energy and data center integration — Div_pradeep · 2026-09-01
- Naver Proposes Verification-Aware Training to Boost Speculative Decoding Draft Models — naver-ai · 2026-09-01
- ExLlamaV3 update: MoE expert CPU offload, GLM-5.3-Flash, self-calibrated quants — Unstable_Llama · 2026-09-01
- Highlander launches: custom GPU kernels serve realtime video at 50% of competitors' cost — Small-Term672 · 2026-09-01
- GPU Debt Investors Assume Zero Residual Value and Distrust Spot Pricing — AccBalanced · 2026-09-01