27B model tops 100 tok/s on a MacBook via M5-tuned quantization and speculative decoding
teortaxesTex · x · 2026-09-09
The norpadon team co-designed quantization and speculative decoding around Apple's M5 neural accelerators, getting a 27B model to run at over 100 tokens/s on a MacBook — for now on M5+ chips only. The latest version also shows significantly improved performance, especially on prefill.
More from Infra
- Desert Ant Labs launches European on-device AI lab with 18 local models — alexcovo_eth · 2026-09-09
- Fickle GPU availability makes on-prem AI compute worth a second look — generativist · 2026-09-09
- Signal65's PINNACLE agentic AI benchmark gets a full GB300 NVL72 rack — ryanshrout · 2026-09-09
- Open-sourcing cimd-proxy: connect MCP clients to any OIDC provider by URL alone — DerMozart · 2026-09-09
- llama.cpp launches llama.app: one-line local LLM install, zero telemetry — ngxson · 2026-09-09
- Dev sarcastically 'thanks' OpenAI for boosting local, privacy-first LLM inference — ngxson · 2026-09-09