AndroidLife: Qwen3.8-27b runs 60 real phone tasks, fails 43% and cooks the chip to 98.2°C
East-Muffin-6472 · reddit · 2026-09-17
The author introduces AndroidLife, a benchmark running AI agents through 60 real tasks on their daily OnePlus phone (no emulator), graded on device state rather than model self-reports. First up: Alibaba's Qwen3.8-27b in text mode scored 56.7% — best text score so far yet still failing 43% — at 29 steps, 6 minutes, and $0.118 per task, with the chip peaking at 98.2°C and 69% battery drained. Easy/medium/hard buckets: 80.8%/52.9%/23.5%. It handles single apps but loops when crossing three apps in a row; three "wins" didn't hold up (a calendar called clash-free with overlapping events, travel times never verified, a name guessed instead of asking). On planted tasks it asked users only on single-fact cases (never multi-ambiguity) and honestly reported 4 of 7 hallucination-control cases. The full corpus spans 530 tasks, 28 days, 31 apps; 11 models planned.
More from Models
- Users complain GPT 5.6 Sol is slow with a context window that fills up fast — jamesbrooksco · 2026-09-17
- 1B-parameter model plays Doom on-device at 150ms latency, training code coming — joemeno · 2026-09-17
- Altman claims internal model beyond 'Astra' can solve problems top mathematicians can't — nikola_mr64990 · 2026-09-17
- OpenAI cracks a math logjam as 25 Fields medalists sign cautionary letter — nordicinst · 2026-09-17
- OpenAI discloses six 'concerning' AI behavior incidents, adds reporting framework — pstAsiatech · 2026-09-17
- "There Is No Moat": Claude Fan Says Rival Model Now Feels Better Overall — deepakns · 2026-09-17