Apple Silicon achieves ANE+GPU dual acceleration, boosting prefill by 50%
bakawolf123 · reddit · 2026-08-20
A developer successfully unlocked dual acceleration using both the ANE and GPU on Apple Silicon for prefill tasks. By sharding parts of the MLP and GDN, this method achieves approximately 50% faster prefill rates for Qwen3.8 27B q4 on M3 Ultra. Real-world testing on an M1 Pro with 32GB RAM showed a 19% speedup (280 -> 334 tokens/s) for a 9B model, though memory usage roughly doubles. The feature is available in the omlx repository via a custom kernel build.
More from Infra
- Data Centers Use 627M Gallons Daily, Far Less Than Cattle or Power Plants — rohanpaul_ai · 2026-08-20
- NVIDIA Releases CUDA-Q Algorithms, Open-Source Primitives for Fault-Tolerant Quantum Computing — tomaszbednarz · 2026-08-20
- Dolphin Network launches P2P inference network for idle GPUs — toptickcrypto · 2026-08-20
- Epoch estimate: OpenAI spends roughly 1.2–2.4% of compute on safety monitoring — sjgadler · 2026-08-20
- Hot take: Cerebras inference could speed up OpenAI's R&D loop 20x — Scobleizer · 2026-08-20
- AI Agents Are Not Microservices: The Need for Durable Execution — rseroter · 2026-08-20