On-device small model selection: Gemma 4 and Qwen3.5 in 2026
lmoroney · x · 2026-08-22
The article analyzes small models (2B-12B parameters) suitable for on-device deployment on phones and laptops. It prioritizes Google's Gemma 4 (Apache 2.0) and Alibaba's Qwen3.5 (Apache 2.0), detailing Gemma 4's size range, performance metrics on Raspberry Pi 5, and the 12B multimodal variant. It compares licensing for SmolLM3 and Llama 4, emphasizing the importance of checking licenses and memory bandwidth for local deployment, and advocates for a hybrid architecture where local models handle simple queries and cloud models handle complex ones.
Related event: Gemma 4 and Qwen3.5 Top Picks for On-Device Models in 2026(2 posts)→
More from Infra
- Qwen3.8-27B gets DFlash2 speculative-decoding GGUF release for llama.cpp — incoai · 2026-08-22
- The Embedder's Dilemma: LLMs match embedding models but cost far more — Adnan El Assadi · 2026-08-22
- Woolly post-trains Qwen3-8B for 2–3× faster math & code decoding — bosmeny · 2026-08-22
- Opinion: States Banning Data Centers Face 20 Years of Economic Depression — GabGarrett · 2026-08-22
- Vercel Fixes TLS Fragmentation Issue, All Websites Back Online — uwukko · 2026-08-22
- Opinion: Data Center Workers and Residents Should Receive Token Dividends — beffjezos · 2026-08-22