Mobile inference spans 30x: LFM fastest at 0.9s, Falcon slowest at 26.7s
ArtificialAnlys · x · 2026-08-25
Real-world tests on iPhone 17 Pro reveal a 30x disparity in end-to-end generation time for 256 output tokens, ranging from 0.9s (LFM2.5-230M) to 26.7s (Falcon-H1R-7B). Top-tier models took 8.0s and 21.4s respectively. The analysis notes that achieving a score of 60 can cost between 5M and 75M output tokens, with Qwen3.5 9B hitting the 16K window limit on 29% of generations, a critical constraint on mobile devices.
Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→
More from Infra
- sPTC speeds up agents via speculative tool calling — a1zhang · 2026-08-25
- Speculative Programmatic Tool Calling Overlaps Code Gen and LLM Inference — a1zhang · 2026-08-25
- Vinci Hits 100k Physics Sims in 24 Hours on Single H200 Node — AnneliesGamble · 2026-08-25
- Cerebras launches CS-4 system with 30x faster inference than GPUs — thione · 2026-08-25
- Local AI Test: Gemma 4 vs Qwen3.8 for Edge and Single-GPU Deployment — lmoroney · 2026-08-25
- Abacus AI Launches SuperComputer: Cloud Workstation with 100+ AI Models — thetripathi58 · 2026-08-25