Run 284B-parameter models locally on a MacBook

socialwithaayan · x · 2026-08-26

Developers demonstrated the ability to run large-scale models locally on a MacBook: DeepSeek-V4-Flash decodes at 5.71 tokens/sec on an M5 Pro with 20GB memory, and a 122B model achieves 16.53 tps. This is enabled by Palm-Infra, which streams experts directly from the SSD, allowing models that llama.cpp cannot even load to run. The project is 100% open source.

Related event: Tencent Open-Sources Palm-Infra, Running 284B Models on a MacBook(2 posts)→

Original post →

More from Infra

Infra channel →