New local AI app hits 164 tokens/sec on iPad, tests to follow

TheOyinbooke · x · 2026-09-26

A developer teased an unreleased local AI app running at 164 tokens/sec on an iPad, promising more benchmarks and opening sign-ups for testers. If confirmed, that on-device speed would make local LLMs notably more usable on mobile, though details on the model, quantization, and hardware are not yet disclosed.

Original post →

More from Infra

Infra channel →