< Back to all clusters
[TECHNOLOGY] · United States · 2 sources

Google revamps Android Bench, puts Claude Fable 5 atop AI coding rankings

Google has overhauled its Android Bench benchmark, replacing the mini‑swe‑agent tool with a new framework called Harbor. The change, which runs tests in secure sandboxes, prompted a complete re‑scoring of all AI models that generate Android code. Claude Fable 5 topped the new leaderboard with a score of 84.5, followed by GPT‑5.5 at 80.2 and Claude Sonnet 5 at 76.2. Google also added eight new models to the list, including Claude Opus 4, GLM 5.2, Kimi K2.7 Code and Qwen 3.7 variants. The company opened the benchmark on GitHub, allowing developers to submit their own Android coding tasks for evaluation, making the platform more transparent and community‑driven.

The updated Android Bench focuses on real‑world Android development scenarios such as migrating code to Jetpack Compose, handling wearable networking, and fixing project‑specific bugs. While Anthropic’s Claude models lead the rankings, Google’s own Gemini models lag behind, highlighting the competitive landscape of AI‑assisted software engineering.