Patching MLX to stage quantized weights to FP8 yields +40% prefill on M6
Brilliant-Hall1387 · reddit · 2026-10-09
A developer found that MLX's quantized matmul (QMM) stages operands to FP16 before multiplying; patching it to stage to FP8 unlocks the 2x faster FP8 matrix path on Apple Silicon M6.
Key points:
- Same trick works on M5: stage 4-bit affine weights to int8 before the matmul
- Gains are prefill-only: +40% for Qwen3-8B, +50% for Qwen3.8-27B; decode stays on FP16 staging since it's bandwidth-bound
- Quality measured: worst-case perplexity increase under 1%
- FP8 needs M6 + macOS 27 with a deployment-target-27 build; int8 works on M5 and later
- Experimental open-source fork with full details, code and evidence on the blog
More from Infra
- Atomic Agent Desktop goes open-source: local Qwen/Gemma agents, cloud planning, 69.8% on GAIA L1 — testingcatalog · 2026-10-09
- YC hosts inference-focused Paper Club: naive vs tuned inference can differ 100x in cost — ycombinator · 2026-10-09
- Cloudflare Workers can now sit next to your PlanetScale database region, edge FUD debunked — ritakozlov · 2026-10-09
- Phinity Exits Stealth With $5.2M Seed for Autonomous Chip Design, 8-Figure ARR — JeffDean · 2026-10-09
- Chollet: AI capex is growing super-exponentially while progress is only sub-linear — fchollet · 2026-10-09
- a16z: Agents burn 5x the tokens of humans, up 14x in six months as AWS rewires for machine users — a16z · 2026-10-09