Hacker Runs 300B Model on M5 Air 32GB: 1tps Decode Using Streamed Experts
maddie-lovelace · reddit · 2026-08-04
A hacker shared an extreme experiment running a 300B parameter model (DeepSeek v4 Flash 4bit) on an M5 Air laptop with 32GB of RAM.
Performance & Optimization Details:
- Speed: By utilizing the 'streamed experts' trick, they achieved 50 tps prefill and 1 tps decode speeds.
- Latency: The initial naive implementation incurred a 30s tax between turns, which was optimized down to 3s.
- Fun Discovery: Running fewer experts than default at prefill time still yields >95% the same top logits at each position. Even when pushed aggressively, performance drops aren't always catastrophic (e.g., needle in haystack performance can be retained).
More from Infra
- ASML Supplier Zeiss Confirms Capability to Meet Surging AI Parts Demand — pstAsiatech · 2026-08-04
- Teaser: Hardware Gauntlet Testing Local LLMs Across PCIe and GPU Configs — TheZachMueller · 2026-08-04
- How Mexico Became a Cornerstone of America's AI Boom via Server Manufacturing — pstAsiatech · 2026-08-04
- Google Backs $200B Infrastructure Financing for Anthropic's AI Chips — firstadopter · 2026-08-04
- Legacy SSE Transport in MCP Causes Serverless Bills to Skyrocket — Ranorkk · 2026-08-04
- RTX 3060 Test: Generates 10-Sec Pixar-Style Animation Locally in 14 Minutes — Pitiful_Archer_4381 · 2026-08-04