Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks
Artificial Analysis, in partnership with Liquid AI, has released benchmarks for small-model intelligence and inference on mobile devices. Multiple models at 4-bit precision and below were tested on real hardware including the iPhone 17 Pro and Galaxy S26 Ultra, with intelligence evaluations capped at a 16K context to simulate phone memory constraints. The results show that intelligence gaps between models are far smaller than efficiency gaps — on-device deployment is bottlenecked more by memory and speed.
Confirmed
- The intelligence evaluation comprises five assessments measuring small models' function calling, knowledge, and reasoning; the inference benchmark measures speed, latency, and memory footprint on real devices.
- In the 16K-context-limited composite evaluation, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tied for first with an average score of 63; Qwen3.5 9B took the lead on BFCL and GPQA Diamond individually.
- The efficiency gap is huge: peak memory at 4K context on the iPhone 17 Pro ranges from 0.4GB (LFM2.5-230M) to 6.9GB (Ornith-1.0-9B and Qwen3.5 9B) — roughly a 19x difference; on a 12GB phone, nearly 7GB of usage severely squeezes system resources.
- End-to-end time to generate 256 tokens ranges from 0.9 seconds (LFM2.5-230M) to 26.7 seconds (Falcon-H1R-7B), about a 30x spread; the two leaders on the performance leaderboard also take noticeably longer.
- Token consumption required to reach 60 points varies widely: Qwen3.5 9B (Reasoning) consumed 74.5M output tokens with 29% of generations exceeding the 16K window; Gemma 4 E4B achieved similar scores with fewer tokens and better efficiency.
Why it matters
- The benchmark systematically quantifies, for the first time, the intelligence-efficiency tradeoff of small models on real hardware: models with similar benchmark scores can differ by an order of magnitude in actual memory usage and latency — far more useful for on-device model selection (especially for mass-market phones with 12GB of RAM).
- The results also highlight the on-device adaptation challenges of long chain-of-thought reasoning models: generations exceeding the context window (such as Qwen3.5 9B's 29%) are truncated outright, hurting usability.
2026-08-25 ~ 2026-08-25 · 8 related posts
Primary sources
- Artificial Analysis launches benchmarks for small models on mobile phones — ArtificialAnlys ·
- Nanbeige and LFM2.5 tie for top mobile model score of 63 — ArtificialAnlys ·
- Mobile memory usage spans 19x: 9B models consume nearly 7GB peak — ArtificialAnlys ·
- Mobile Benchmark: LFM2.5 Leads Efficiency on iPhone 17 Pro — ArtificialAnlys · 2026-08-25
- [source] Nanbeige and LFM2.5 tie for top mobile model score of 63 — ArtificialAnlys · 2026-08-25
- Qwen3.5 consumes 14x more tokens than Gemma 4 for similar mobile scores — ArtificialAnlys · 2026-08-25
- Mobile inference spans 30x: LFM fastest at 0.9s, Falcon slowest at 26.7s — ArtificialAnlys · 2026-08-25
- [source] Mobile memory usage spans 19x: 9B models consume nearly 7GB peak — ArtificialAnlys · 2026-08-25
- [source] Artificial Analysis launches benchmarks for small models on mobile phones — ArtificialAnlys · 2026-08-25
- Artificial Analysis partners with Liquid for mobile benchmarks — JosephJacks_ · 2026-08-25
- LiquidAI Partners for Mobile Small Model Benchmarking — JosephJacks_ · 2026-08-25