Gemma 4 26B runs on an iPhone 17 Pro via model paging
Agreeable-Rest9162 · reddit · 2026-07-25
Noema founder says Gemma 4 26B A4B is running on an iPhone 17 Pro using model paging.
The setup keeps non-expert weights in RAM and reads expert weights from SSD, trading speed for the ability to run a much larger model on-device. The post reports a 699-token prompt, 34.4 tokens/s prefill, 20.34 seconds prefill time, and 3.5 tokens/s decode speed. The full answer took about 6 minutes, but the author argues that the approach could be useful when accuracy matters more than latency, including on low-RAM MacBooks.
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27