tech
Google updates Android Bench with new LLMs, but Gemini still lags behind
Android Bench is evolving, and developers can help guide that process.

TL;DR
- Google has updated its Android Bench benchmark with eight new LLMs and a new framework called Harbor.
- Claude Fable 5 leads in accuracy with 84.5%, but has high operating costs.
- Google's Gemini 3.1 Pro is in fifth place, with lower accuracy but also lower costs than top performers.
- Gemini 3.5 Flash has the highest cost due to a long runtime.
- Google is encouraging developers to contribute to Android Bench through the Harbor framework.