Turning an iPhone into a second GPU for a MacBook: 44% faster prefill on Qwen 27B
StayLameBro · reddit · 2026-10-03
A developer wired an iPhone 17 Pro Max to a 24GB M4 Pro MacBook over a 10Gb/s USB-C cable so the phone's A19 Pro GPU runs layers 41–64 of Qwen 3.8 27B (IQ4XS) in a pipeline with the Mac: prefill speeds up 29–44% depending on context length, and a 27k-token agent session cold-start drops from 245s to 168s. Past 64k context the phone instead holds old KV pages and computes attention over them, extending 8-bit context to 196k–229k with up to 5.7GB of KV cache offloaded; at 140k context, writing drops from 279 to 176 ms/token. Their fork's SME2 kernels, Metal fusions and DFlash2 speculative decoding also double writing speed to 25 tok/s. Code and benchmarks are open-sourced (backburner on GitHub), and the project was built with Claude Opus 5.5.
More from coding & agent
- jasonkneen: every personal agent has the same unsolved problem — jasonkneen · 2026-10-03
- Dev builds WoW private server, browser client, and an MCP agent to play it — ProfessorMunchy · 2026-10-03
- Y Combinator's QM team-oriented AI agent harness is now free on Agent37 — ycombinator · 2026-10-03
- OpenAI's Dot Agent Spotted Production Outage Five Minutes Before DevDay Live Demo — lennysan · 2026-10-03
- LangChain's Chase Amps Up 'Durable Harness' Debate as Pi Durable Draws HN Praise — hwchase17 · 2026-10-03
- OpenAI Dot's Codex task access only works about 50% of the time, user reports — Danny7Guns · 2026-10-03