Researchers from LMSYS have released a leaderboard using an Elo rating system to compare large language models, revealing that Anthropic’s Claude 3 Opus has secured the top spot. The model achieved an Elo rating of 1253 based on head-to-head battles evaluated by human voters.

  • Claude 3 Opus holds the highest rating with 33,250 total votes and a +5/-5 record.
  • OpenAI’s GPT-4 preview models follow closely with ratings of 1251 and 1248.
  • Google’s Bard (Gemini Pro) ranks fourth with an Elo rating of 1203.
  • The older GPT-4–0314 model is ranked sixth with a rating of 1185.

This benchmark provides an objective method to track the rapid advancements in language AI and identify leading architectures as competition intensifies among major tech companies.