Apple Silicon achieves ANE+GPU dual acceleration, boosting prefill by 50%

bakawolf123 · reddit · 2026-08-20

A developer successfully unlocked dual acceleration using both the ANE and GPU on Apple Silicon for prefill tasks. By sharding parts of the MLP and GDN, this method achieves approximately 50% faster prefill rates for Qwen3.8 27B q4 on M3 Ultra. Real-world testing on an M1 Pro with 32GB RAM showed a 19% speedup (280 -> 334 tokens/s) for a 9B model, though memory usage roughly doubles. The feature is available in the omlx repository via a custom kernel build.

Original post →

More from Infra

Infra channel →