Qwen3.5 consumes 14x more tokens than Gemma 4 for similar mobile scores

ArtificialAnlys · x · 2026-08-25

Achieving a score of 60 can cost 5M or 75M output tokens depending on the model. Qwen3.5 9B (Reasoning) spends 74.5M tokens and exceeds the 16K window on 29% of generations, while Gemma 4 E4B (Reasoning) reaches a similar score on just 5.2M tokens without hitting the limit. On mobile phones, token usage translates directly into time, energy, and heat, which are less tolerable than on servers.

Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→

Original post →

More from Models

Models channel →