Qwen3.5 consumes 14x more tokens than Gemma 4 for similar mobile scores
ArtificialAnlys · x · 2026-08-25
Achieving a score of 60 can cost 5M or 75M output tokens depending on the model. Qwen3.5 9B (Reasoning) spends 74.5M tokens and exceeds the 16K window on 29% of generations, while Gemma 4 E4B (Reasoning) reaches a similar score on just 5.2M tokens without hitting the limit. On mobile phones, token usage translates directly into time, energy, and heat, which are less tolerable than on servers.
Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→
More from Models
- Polymarket Forecast on DeepSeek Pro; Bloomberg Warns of Hacker Use — Polymarket · 2026-08-25
- ComfyUI nodes enable grafting Krea2 features into MiniMax H3 — Key-Philosopher-9327 · 2026-08-25
- Planning $100 benchmark for Qwen quantization and KV cache trade-offs — m_mukhtar · 2026-08-25
- Users Frustrated: LLMs Cite Scholarly Consensus to Dismiss Your Own Claims — louisvarge · 2026-08-25
- OpenCode Go Offers $60 of Usage for a $10 Monthly Sub — chrisalbon · 2026-08-25
- GPT-5.6 Sol Max beats Fable 5 Max on DeepSWE at $6.47 vs $21.63 per task — reach_vb · 2026-08-25