Patching MLX to stage quantized weights to FP8 yields +40% prefill on M6

Brilliant-Hall1387 · reddit · 2026-10-09

A developer found that MLX's quantized matmul (QMM) stages operands to FP16 before multiplying; patching it to stage to FP8 unlocks the 2x faster FP8 matrix path on Apple Silicon M6.

Key points:

Original post →

More from Infra

Infra channel →