Solo researcher boosts Qwen 3.8 on Mac: +54% decode at 147k context
julianharris · x · 2026-08-29
Independent researcher Youssof released MTPLX V2.10, a full-stack performance overhaul for running Qwen 3.8 on Mac — praised by Julian Harris as a one-man local AI powerhouse, spanning low-level kernel optimizations inside the model plus a custom inference server. Highlights:
- Qwen 3.8 Next support: 4-bit and optimized 4/8-bit models sustain 50–80 TPS at high contexts
- mmap streaming: the 32 GB n-gram table streams from SSD via mmap instead of RAM; both models fit a 96 GB Mac with zero speed impact
- 27B decode & prefill boost everywhere: +15% TPS at 3k context, +29% at 88k, +54% at 147k; prefill +41%
- 48 GB Macs: from 3–4 TPS in swap at 30k context up to 33 TPS
- KV cache solved: 8-bit and 4-bit caches now barely hurt speed, versus a prior 50% decline
Harris suggests Jeremy Howard's Answer.ai should snap him up before Big Tech does.
More from Infra
- Qwen3.8 IQ1_S on RTX 5070: Achieving 22t/s with 12GB VRAM — jacek2023 · 2026-08-29
- Exo Labs claims 4.8 TB/s memory bandwidth from clustered Mac Studios, scaling linearly — anonmt57 · 2026-08-29
- Lightning AI deploys H200s and launches VMs early access — LightningAI · 2026-08-29
- Open Source Personal AI Datacenter with Full Guides — dee_hw · 2026-08-29
- Nvidia's AI edge is moving beyond the GPU with smarter data center traffic control — TechCrunch AI · 2026-08-29
- One Beef Burger Equals Lifetime of ChatGPT Use in Water Usage — dc_lawrence · 2026-08-29