AndroidLife real-phone benchmark: Qwen3.8-27B fails 43% of 60 daily tasks

East-Muffin-6472 · reddit · 2026-09-18

AndroidLife is a new benchmark running AI agents on a real OnePlus phone—60 public tasks from a 530-task corpus across 31 apps, graded on actual device state. First run: Alibaba's Qwen3.8-27B in text mode scored 56.7%, the best text-model score yet, still failing 43%. Per task: 29.25 steps, 6 minutes, $0.118, 69% battery drain, 98.2°C peak chip temp. Easy/medium/hard buckets: 80.8%/52.9%/23.5%; single apps mostly work but three-app sequences exhaust the step limit. Three claimed wins didn't hold on-device. On planted tasks, it asked the user only on single-fact questions, never multi-question ones, and honestly reported missing data in 4 of 7 hallucination controls. Leaderboard and trajectories public; 10 more models to come.

Original post →

More from Models

Models channel →