Google expands Android Bench with eight new LLMs, Gemini falls to fifth place
Google has upgraded its Android Bench benchmark, which evaluates large language models (LLMs) on 100 Android development tasks. The update adds eight new models—Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max—and introduces metrics for cost and efficiency as well as support for open‑weight models. Developers are invited to run their own tests and submit feedback.
In the revised leaderboard, Google’s Gemini 3.1 Pro ranks fifth, trailing OpenAI’s GPT 5.4, Claude Sonnet 5 and Claude Fable 5. Claude Fable 5 leads with 84.5% accuracy, reinforcing its strong performance in Android code generation tasks.
The benchmark aims to help developers identify the most capable LLMs for Android app development and to guide future improvements to the models and the benchmark itself.