AndroidLife: Qwen3.8-27b runs 60 real phone tasks, fails 43% and cooks the chip to 98.2°C

East-Muffin-6472 · reddit · 2026-09-17

The author introduces AndroidLife, a benchmark running AI agents through 60 real tasks on their daily OnePlus phone (no emulator), graded on device state rather than model self-reports. First up: Alibaba's Qwen3.8-27b in text mode scored 56.7% — best text score so far yet still failing 43% — at 29 steps, 6 minutes, and $0.118 per task, with the chip peaking at 98.2°C and 69% battery drained. Easy/medium/hard buckets: 80.8%/52.9%/23.5%. It handles single apps but loops when crossing three apps in a row; three "wins" didn't hold up (a calendar called clash-free with overlapping events, travel times never verified, a name guessed instead of asking). On planted tasks it asked users only on single-fact cases (never multi-ambiguity) and honestly reported 4 of 7 hallucination-control cases. The full corpus spans 530 tasks, 28 days, 31 apps; 11 models planned.

Original post →

More from Models

Models channel →