AndroidLife real-phone benchmark: Qwen3.8-27B fails 43% of 60 daily tasks
East-Muffin-6472 · reddit · 2026-09-18
AndroidLife is a new benchmark running AI agents on a real OnePlus phone—60 public tasks from a 530-task corpus across 31 apps, graded on actual device state. First run: Alibaba's Qwen3.8-27B in text mode scored 56.7%, the best text-model score yet, still failing 43%. Per task: 29.25 steps, 6 minutes, $0.118, 69% battery drain, 98.2°C peak chip temp. Easy/medium/hard buckets: 80.8%/52.9%/23.5%; single apps mostly work but three-app sequences exhaust the step limit. Three claimed wins didn't hold on-device. On planted tasks, it asked the user only on single-fact questions, never multi-question ones, and honestly reported missing data in 4 of 7 hallucination controls. Leaderboard and trajectories public; 10 more models to come.
More from Models
- natolambert: latest OpenAI jailbreak via Claude shows closed models are the real AI risk — natolambert · 2026-09-18
- Open-source models already at SOTA — Anthropic/OpenAI edge is just 5GW compute, dev argues — ccerrato147 · 2026-09-18
- Grok Voice Transcribe 2.0 hits 92.9% accuracy on phone calls at just $0.10/hour — XFreeze · 2026-09-18
- Meta's 'Muse Spark' spotted on Hugging Face, release appears imminent — eliebakouch · 2026-09-18
- Silent Model Revisions Break Everything — Versioning and Model Cards Are a Mess — xeophon · 2026-09-18
- Suspected Gemini 4 Pro generates animated peacock-on-a-bike SVG in Arena — CodeByPoonam · 2026-09-18