On-device Android agent with Gemma 4 E2B hits 2.6 tok/s live vs 11 tok/s on replay
HowDevelop · x · 2026-09-07
Developer HowDevelop is building an on-device Android agent using Gemma 4 E2B with Qualcomm QMX kernels. Side note: he spent 5+ hours debugging with GPT-6 Astra on medium, burning 40% of his usage without a confirmed root cause.
The key puzzle: the live agent runs at 2.6 tok/s, while replaying the same request hits 11 tok/s—the gap is unexplained so far.
Related event: Developer Tests On-Device Android Agent with Gemma 4 E2B(2 posts)→
More from coding & agent
- LLM-powered revival of Put-That-There brings speech and gesture window control to XR — twi_mar · 2026-09-07
- Anthropic Claude Code engineer on internal AI coding practices and autonomy share — trq212 · 2026-09-07
- Leak: OpenAI to unveil Managed Agents at DevDay 2026 with hosted or self-hosted deploy — testingcatalog · 2026-09-07
- Harness-only changes lift deepagents-cli from 52.8% to 66.5% on Terminal-Bench 2.0 — Gauri_the_great · 2026-09-07
- Chops: open-source macOS app to manage AI agent skills across Claude Code, Cursor, Codex — tom_doerr · 2026-09-07
- LLMs vs classic CV: 200-line pipelines beat token-burning agents on rote tasks — mervenoyann · 2026-09-07