Turning an iPhone into a second GPU for a MacBook: 44% faster prefill on Qwen 27B

StayLameBro · reddit · 2026-10-03

A developer wired an iPhone 17 Pro Max to a 24GB M4 Pro MacBook over a 10Gb/s USB-C cable so the phone's A19 Pro GPU runs layers 41–64 of Qwen 3.8 27B (IQ4XS) in a pipeline with the Mac: prefill speeds up 29–44% depending on context length, and a 27k-token agent session cold-start drops from 245s to 168s. Past 64k context the phone instead holds old KV pages and computes attention over them, extending 8-bit context to 196k–229k with up to 5.7GB of KV cache offloaded; at 140k context, writing drops from 279 to 176 ms/token. Their fork's SME2 kernels, Metal fusions and DFlash2 speculative decoding also double writing speed to 25 tok/s. Code and benchmarks are open-sourced (backburner on GitHub), and the project was built with Claude Opus 5.5.

Original post →

More from coding & agent

coding & agent channel →