Palm-Infra open-sourced: Run 284B models on MacBook via SSD offloading
dr_cintas · x · 2026-08-26
Tencent's Palm-Infra framework enables running massive MoE models on Apple Silicon by streaming experts directly from SSD. Benchmarks show DeepSeek-V4-Flash (284B) decoding at 5.71 tps and a 122B model at 16.53 tps on an M5 Pro with 20GB memory, achieving capabilities previously impossible with llama.cpp.
Related event: Tencent Open-Sources Palm-Infra, Running 284B Models on a MacBook(2 posts)→
More from Infra
- Glean reveals model routing scores: GPT-5.6 Luna leads at $0.08 — testingcatalog · 2026-08-27
- Serving frontier models at scale on purely Chinese hardware — tokumin · 2026-08-27
- Firecrawl Launches Startup Deal: Up to $30k in Credits — devdigest · 2026-08-27
- Long Read: AI Is Buying the Data of Dead Companies — rvp · 2026-08-27
- Antirez: High prefill speed makes LLMs feel 10x more powerful — antirez · 2026-08-27
- Beijing startup Bolun Zhihui raises angel funding for heterogeneous GPU scheduling architecture — 智东西 · 2026-08-27