AndroidLife runs Qwen3.8-27b on a real phone: 56.7% success over 60 tasks, 98.2°C chip, 69% battery

East-Muffin-6472 · reddit · 2026-09-18

AndroidLife is a benchmark that runs AI agents on a real Android phone: 60 public tasks (from a 530-task corpus, 28 days, 31 everyday apps), graded on device state — never the model's own report — with thermal, battery and cost telemetry. First run: Alibaba's qwen3.8-27b in text mode.

Key numbers:

Full leaderboard, tasks and trajectories are public. Long-horizon multi-app operation and proactive user asking remain the biggest real-device gaps.

Original post →

More from coding & agent

coding & agent channel →