AndroidLife runs Qwen3.8-27b on a real phone: 56.7% success over 60 tasks, 98.2°C chip, 69% battery
East-Muffin-6472 · reddit · 2026-09-18
AndroidLife is a benchmark that runs AI agents on a real Android phone: 60 public tasks (from a 530-task corpus, 28 days, 31 everyday apps), graded on device state — never the model's own report — with thermal, battery and cost telemetry. First run: Alibaba's qwen3.8-27b in text mode.
Key numbers:
- 56.7% success (best text score on the board, still failing 43% of 60 tasks)
- 29.25 steps, 6 min and $0.118 per task; 69% battery drained, chip peaked at 98.2°C, skin 48.9°C
- Difficulty buckets: easy 80.8%, medium 52.9%, hard 23.5%
- Single-app tasks mostly work; three apps in a row means tapping in circles until the step limit — most of the 43%
- 3 claimed wins didn't hold up on-device (calendar clash misjudged, travel times never opened, a name guessed instead of asking)
- Planted probes: asked the user only twice across 11 ASK USER tasks (never on MULTI ones); honestly reported missing data on 4 of 7 hallucination controls
Full leaderboard, tasks and trajectories are public. Long-horizon multi-app operation and proactive user asking remain the biggest real-device gaps.
More from coding & agent
- Anthropic's Head of Product Drops a 28-Minute Masterclass on Agents in Production — ifioknkem · 2026-09-20
- Teknium: Jev can't compact context well — Hermes summarizes 95% of it away — Teknium · 2026-09-20
- HarnessRouter: routing agent harnesses instead of models, a fresh infra idea — daniel_mac8 · 2026-09-20
- MCP tool naming: short generic verbs vs explicit prefixes for LLM tool selection — skvark · 2026-09-20
- GameToMac launched 10 days ago and already runs AoE IV, CS2 and Diablo IV on Apple Silicon — nickbaumann_ · 2026-09-20
- DialKit 2.0 ships: open-source real-time UI tuning tool with prompts for coding agents — LinusEkenstam · 2026-09-20