Laurence Moroney on 2026 On-Device Small AI: Gemma 4 & Qwen 3.5 Top Picks
lmoroney · x · 2026-08-22
Laurence Moroney analyzes the state of small models (2-12B parameters) suitable for on-device deployment in 2026, focusing on permissive licenses and performance for phones/laptops/offline use.
- Gemma 4:
- License: Switched to Apache 2.0, making it ideal for education and commercial use.
- Specs: Ranges from 2.3B (E2B) to 31B. The new 12B Unified is an encoder-free multimodal model running on 16GB VRAM with native audio.
- Performance: E2B achieves 133 prefill/7.6 decode tok/s on Raspberry Pi 5 CPU, higher on NPUs.
- Qwen 3.5: Also Apache 2.0, another pillar of the on-device ecosystem.
- Others: SmolLM3 publishes a full recipe; Llama 4 is open-weight but requires license review.
Moroney recommends prioritizing Gemma 4 and Qwen 3.5 for local stacks, emphasizing memory bandwidth and reproducibility.
Related event: Gemma 4 and Qwen3.5 Top Picks for On-Device Models in 2026(2 posts)→
More from Infra
- MCP Isn't Replacing APIs: It's Changing Who APIs Are Designed For — kush_patil · 2026-08-22
- Data Center Opposition Surged from 42 to 75 Percent in One Year — The Decoder · 2026-08-22
- Qwen3.8-27B gets DFlash2 speculative-decoding GGUF release for llama.cpp — incoai · 2026-08-22
- Woolly post-trains Qwen3-8B for 2–3× faster math & code decoding — bosmeny · 2026-08-22
- Opinion: States Banning Data Centers Face 20 Years of Economic Depression — GabGarrett · 2026-08-22
- Vercel Fixes TLS Fragmentation Issue, All Websites Back Online — uwukko · 2026-08-22