Running AI Agents on CPU: Gemma 4 26B vs Qwen 3.6 35B?
Kahvana · reddit · 2026-08-18
A user asks for advice on choosing a model to run an Agent on CPU for web search and document parsing, requiring 256K context support. Options include Gemma 4 26B-A4B (for natural language skills) and Qwen 3.6 35B-A3B (for lighter resource usage). The user has 96GB RAM but GPU VRAM is fully occupied. Seeking recommendations for CPU-friendly agent models and experiences with what works or doesn't.
More from coding & agent
- Permix: Lightweight Type-Safe TS Permissions Library Wins Developer Praise — jonathan_wilke · 2026-08-18
- Give agents a centralized source of truth before they hallucinate — eptwts · 2026-08-18
- After letting her agents queue up commits, her GitHub daily streak exploded to 2026 — christine_hall · 2026-08-18
- Why logging is vital for agent apps: accountability, evals, and self-healing — jasonkneen · 2026-08-18
- Supabase ships battle-tested UI components — hand them to your agent — dshukertjr · 2026-08-18
- PyLate hands model maintenance to SentenceTransformers' Tom Aarsen — antoine_chaffin · 2026-08-18