Google's Android Bench 2.0: GPT 6 Astra Leads Long-Horizon Android Coding at 28% Pass Rate
sandersted · x · 2026-09-23
Google launched Android Bench 2.0, adding long-horizon agent results on 30 real-world Android app development tasks. GPT 6 Astra (codex) leads at a 28.0% pass rate ($375.7/run), ahead of Claude Fable 5.1 (22.7%, $492.6) and GPT 5.6 Sol (19.3%); Claude Opus 5 scores 16.7% at a hefty $861.4 per run, while Gemini 3.8 Flash trails on pass rate but is among the fastest and cheapest.
More from coding & agent
- Half of PM job listings now ask for evals skills; Ramp, Shopify, Harvey, Cursor share gains — HamelHusain · 2026-09-23
- Security audit: autonomous research program XBOW credited with ~12 upstream bug fixes — moyix · 2026-09-23
- Telling agents to use formal verification helps them write better code — sh_reya · 2026-09-23
- Team uses Typesafe's Jev to prompt follow-up questions during ticket creation — TheMoonMidas · 2026-09-23
- The most common evals mistake: skipping error discovery and measuring the wrong thing — FinanceYF5 · 2026-09-23
- AI products are easy to change and hard to predict: evals turn 'good' into repeatable tests — FinanceYF5 · 2026-09-23