Run 284B-parameter models locally on a MacBook
socialwithaayan · x · 2026-08-26
Developers demonstrated the ability to run large-scale models locally on a MacBook: DeepSeek-V4-Flash decodes at 5.71 tokens/sec on an M5 Pro with 20GB memory, and a 122B model achieves 16.53 tps. This is enabled by Palm-Infra, which streams experts directly from the SSD, allowing models that llama.cpp cannot even load to run. The project is 100% open source.
Related event: Tencent Open-Sources Palm-Infra, Running 284B Models on a MacBook(2 posts)→
More from Infra
- Turso Adopts AgentID to Grant AI Agents Independent OIDC Identities — glcst · 2026-08-27
- Sail Research CEO on building extreme-efficiency inference infra for long-running agents — agihouse_org · 2026-08-27
- Deep Dive into Zhipu GLM-5.3-Flash: Architecture Overhaul and Domestic Infrastructure Breakthrough — 赛博禅心 · 2026-08-26
- Idea: Blockchain-based prompt credentialing for AI models — Dsphar · 2026-08-26
- Foresight CEO: Open Science Needs Independent Secure Compute Clusters — allisondman · 2026-08-26
- Qwen3.8-Flash Runs Locally: 125B Model on Just 75GB RAM — danielhanchen · 2026-08-26