What MoE/LLM runs well offline on a 24GB M5 MacBook Air?
itis_whatit-is · reddit · 2026-09-12
A Reddit user asks which MoE or LLM is fast and good enough for basic chat on a 24GB M5 MacBook Air in a fully offline scenario. They already run an AI home server (16GB/64GB) but want a practical recommendation for on-the-go offline use.
More from Infra
- Glass core substrates show 2x better warpage than organic core without stiffener — jwt0625 · 2026-09-12
- d-Matrix partners with NVIDIA to plug Raptor XPUs into NVLink Fusion rackscale systems — bookwormengr · 2026-09-12
- Running out of context on a large codebase: how to auto-handoff long-running local LLM tasks — Developer-Y · 2026-09-12
- The Economist: Nvidia is the central bank of AI — tolugenius · 2026-09-12
- Draft model hits ~60 tok/s running Qwen3.8-27B at 131k context on a 16GB GPU — pneuny · 2026-09-12
- DeepSeek v4.1 Flash on-device test: q2 runs at 16 tok/s but tool calls go off the rails — challis88ocarina · 2026-09-12