The AA benchmark has been updated to reflect the latest performance of frontier language models. The update includes fresh testing data for top-tier models.

  • Flash-next is reported to outperform 3.5 max in the new rankings.
  • Gemini 3.1 pro is noted as being significantly behind other contenders.

The article highlights these shifts in the competitive landscape among leading AI models.