Qwen 3.8 in 4-bit hits 25 tok/s on an M3 Max — fully usable locally
AIandDesign · x · 2026-08-19
A user reports running Qwen 3.8 locally with 4-bit quantization on an M3 Max at roughly 25 tokens/sec. It's expected to get faster, but it's already totally usable — the author says it's effectively better than what GPT-4 offered, all on their own machine. "WILD."
More from Infra
- Etched poaches top Nvidia engineers, runs inference in 44 days — SumitGup · 2026-08-19
- Open Source Personal AI Computer: Local Compute Blueprint — tom_doerr · 2026-08-19
- UC Berkeley Releases FreeToken for Efficient Edge-Native MoE Serving — UCBerkeley · 2026-08-19
- Cerebras Unveils New AI Inference System for Faster Chatbot Responses — Polymarket · 2026-08-19
- Podcast: OpenRoboto on training humanoid robot AI on decentralized networks — markjeffrey · 2026-08-19
- GitHub Outage Inspires Durable Objects-based Git Forge — tobowers · 2026-08-19