Gemma 4 26B runs on an iPhone 17 Pro via model paging

Agreeable-Rest9162 · reddit · 2026-07-25

Noema founder says Gemma 4 26B A4B is running on an iPhone 17 Pro using model paging.

The setup keeps non-expert weights in RAM and reads expert weights from SSD, trading speed for the ability to run a much larger model on-device. The post reports a 699-token prompt, 34.4 tokens/s prefill, 20.34 seconds prefill time, and 3.5 tokens/s decode speed. The full answer took about 6 minutes, but the author argues that the approach could be useful when accuracy matters more than latency, including on low-RAM MacBooks.

Original post →

More from Infra

Infra channel →