Android Bench Updates Model Leaderboard

Ars Technica AI · rss · 2026-07-09

Google updated Android Bench, its benchmark for evaluating Android app development capabilities, adding several new models and evaluation dimensions.

This update incorporates cost and efficiency metrics alongside open-weight models, helping developers comprehensively compare LLM performance on Android tasks. Google also invites developers to run tests and provide feedback to shape the benchmark's future.

Newly added models include: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, Qwen 3.7 Max. The article concludes that Gemini still trails some competing models on this leaderboard.

Original post →

More from Models

Models channel →