M5 Macs can go faster: w8a8 kernels lift Gemma4 prefill from 2,193 to 3,029 tps
maddie-lovelace · reddit · 2026-07-24
A Reddit user says Apple M5 silicon is not yet being fully used by current inference backends.
- MLX and llama.cpp still run 16-bit activations everywhere, even though M5 reportedly supports INT8 activations / w4a8 dtype.
- The author built w8a8 kernels and reports 1.4x faster Gemma4 prefill performance.
- On an M5 MacBook Air, prefill throughput rose from 2,193 tps to 3,029 tps on a 130,173-token input.
- For shorter contexts, the author says it gets even faster and approaches 10k tps.
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11