Mobile inference spans 30x: LFM fastest at 0.9s, Falcon slowest at 26.7s

ArtificialAnlys · x · 2026-08-25

Real-world tests on iPhone 17 Pro reveal a 30x disparity in end-to-end generation time for 256 output tokens, ranging from 0.9s (LFM2.5-230M) to 26.7s (Falcon-H1R-7B). Top-tier models took 8.0s and 21.4s respectively. The analysis notes that achieving a score of 60 can cost between 5M and 75M output tokens, with Qwen3.5 9B hitting the 16K window limit on 29% of generations, a critical constraint on mobile devices.

Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→

Original post →

More from Infra

Infra channel →