Qwen 3.8 Flash Next tested on a 128GB Mac: smarter and faster than 27B, needs ~80GB RAM
julianharris · x · 2026-09-01
Julian Harris tested Qwen 3.8 on a MacBook Pro 14" M5 Max / 128GB using MTPLX, a vertical local-LLM stack for Macs:
- 27B throughput drops significantly as the context window grows;
- Flash Next ("Qwen 4 preview", MoE 125-a6b + 4-bit + ngram caching) degrades far less — not as good as cloud Opus, but smarter, faster and more efficient overall;
- That config needs 80GB RAM, so it's really only viable on 128GB Macs; flash-next does clever offloading but performance 'sinks like a stone' AFAICT.
More details coming on his blog AI Local.
Related event: Qwen 3.8 Tested on 128GB Mac: Smarter but Memory-Hungry(2 posts)→
More from Infra
- SmolCluster: Distributed Training Library for Heterogeneous Hardware Using Raw Sockets — retr0jirachi · 2026-09-01
- Milvus: Cloud-Native High-Performance Vector Database — goyalshaliniuk · 2026-09-01
- Weaviate: Vector Database with Structured Filtering — goyalshaliniuk · 2026-09-01
- Saudi AI firms Humain, DataVolt partner on 100MW Red Sea data center — Polymarket · 2026-09-01
- Mixed old GPUs + 32GB RAM runs Qwen3 27B at 20t/s for local coding — pepijndevos · 2026-09-01
- ByteDance's UBASE: AI Search Engine for Trillion-Scale Vectors — _reachsumit · 2026-09-01