Benchmark · general

LMSYS Arena (Elo)

23 results 21 models

LMSYS Chatbot Arena measures how good AI chatbots are at open-ended conversation by having real people compare two anonymous models side by side. Results are aggregated into an Elo rating, the same relative-skill number used in chess (roughly 1000-1500+).

Read more
Example
A visitor types a free-form prompt — say, 'Explain quantum entanglement to a 10-year-old' or 'Write a Python function to reverse a linked list' — sees two hidden models answer, then picks the better reply.
Scoring
Each blind vote is a pairwise win/loss/tie; these votes are pooled across many users and converted into an Elo score, where a higher number means the model wins more of its matchups.
Verification
There are no automated tests or reference answers — a result is decided purely by crowdsourced human preference votes, cast blind so brand names can't bias the choice.
Why it matters
Because it reflects real users' subjective judgment on everyday prompts rather than a fixed test set, it's a widely watched, hard-to-game leaderboard for overall chatbot quality.
Worked example
Task
Chatbot Arena prompt (open-ended, two models compared side-by-side): 'Explain the difference between TCP and UDP to a curious 10-year-old using a simple analogy, then name one everyday app that relies on each.' Two anonymous models answer the same prompt and you vote for the better reply.
Solution
Strong reply: 'TCP = numbered letters the post office confirms and resends if lost, so everything arrives in order — web browsing, email. UDP = shouting across a playground: fast, but missed words are never resent — video calls, online games.' No single gold answer exists; the winner is whichever reply a human rates more helpful, accurate, and clear.
Walkthrough
The reply is judged good because it is technically accurate (TCP = reliable and ordered, UDP = fast, best-effort), uses one age-appropriate analogy, and covers both requested parts (an app for each). Grading: there is no gold label — two anonymous models answer the same prompt, a human picks the better response, and pairwise votes aggregate into an Elo/Bradley-Terry leaderboard.
1000 1250 1500 1750 2000 2024-03-27 2025-05-25 2026-07-24 Bard (Gemini Pro) · 1203 · 2024-03-27 GPT-4 preview models · 1251 · 2024-03-27 GPT-4–0314 · 1185 · 2024-03-27 Claude 3 Sonnet · 1198 · 2024-03-27 Qwen1.5–72B-Chat · 1148 · 2024-03-27 Mistral-Large-2402 · 1157 · 2024-03-27 Claude 3 Opus · 1253 · 2024-03-27 GPT-4o · 1309 · 2024-05-15 Gemini 1.5 Pro 0801 · 1300 · 2024-08-02 Gemini-Exp-1114 · 1344 · 2024-11-15 Gemini-Exp-1206 · 1379 · 2024-12-09 Grok-3 · 1402 · 2025-02-18 Grok 3 · 1402 · 2025-02-19 Gemini 2.5 Pro · 1415 · 2025-05-20 Gemini 2.5 Pro · 1470 · 2025-06-05 Gemini 2.5 06-05 · 1470 · 2025-06-05 Grok4.1 (Thinking) · 1510 · 2025-11-17 Gemini 3 Pro · 1501 · 2025-11-18 Gemini 3 Pro · 1501 · 2025-11-18 Claude Opus 4.6 · 1606 · 2026-02-05 Kimi K3 · 1679 · 2026-07-24 Claude Fable 5 · 1760 · 2026-07-24 GPT-5.6 Sol · 1743 · 2026-07-24
Bard (Gemini Pro) GPT-4 preview models GPT-4–0314 Claude 3 Sonnet Qwen1.5–72B-Chat Mistral-Large-2402 Claude 3 Opus GPT-4o Gemini 1.5 Pro 0801 Gemini-Exp-1114 Gemini-Exp-1206 Grok-3 Grok 3 Gemini 2.5 Pro Gemini 2.5 06-05 Grok4.1 (Thinking) Gemini 3 Pro Claude Opus 4.6 Kimi K3 Claude Fable 5 GPT-5.6 Sol
Timeline
Date Model Score Source
2026-07-24 Kimi K3 1679.0Elo Kimi K3 launches as largest open-weight model; Muse Spark 1.1 released
2026-07-24 Claude Fable 5 1760.0Elo Kimi K3 launches as largest open-weight model; Muse Spark 1.1 released
2026-07-24 GPT-5.6 Sol 1743.0Elo Kimi K3 launches as largest open-weight model; Muse Spark 1.1 released
2026-02-05 Claude Opus 4.6 1606.0Elo Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
2025-11-18 Gemini 3 Pro 1501.0Elo Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Pro 1501.0Elo Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-17 Grok4.1 (Thinking) 1510.0Elo xAI releases Grok 4.1, topping LMArena leaderboard with reduced hallucinations
2025-06-05 Gemini 2.5 06-05 1470.0Elo Google releases Gemini 2.5 06-05 preview with coding and reasoning upgrades
2025-06-05 Gemini 2.5 Pro 1470.0Elo Google introduces upgraded Gemini 2.5 Pro preview with improved coding and reasoning
2025-05-20 Gemini 2.5 Pro 1415.0Elo Google DeepMind updates Gemini 2.5 with Deep Think, audio output, and computer use
2025-02-19 Grok 3 1402.0Elo xAI releases Grok 3 Beta and DeepSearch agent
2025-02-18 Grok-3 1402.0Elo xAI's Grok-3 becomes first model to surpass 1400 in Chatbot Arena
2024-12-09 Gemini-Exp-1206 1379.0Elo Google's Gemini-Exp-1206 surpasses ChatGPT on LMArena leaderboard
2024-11-15 Gemini-Exp-1114 1344.0Elo Google's Gemini-Exp-1114 ranks joint No. 1 on Chatbot Arena
2024-08-02 Gemini 1.5 Pro 0801 1300.0Elo Google's experimental Gemini 1.5 Pro 0801 surpasses GPT-4o on LMSYS Chatbot Arena
2024-05-15 GPT-4o 1309.0Elo OpenAI secretly tested GPT-4o in Chatbot Arena, topping the leaderboard
2024-03-27 Bard (Gemini Pro) 1203.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 GPT-4 preview models 1251.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 GPT-4–0314 1185.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 Claude 3 Sonnet 1198.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 Qwen1.5–72B-Chat 1148.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 Mistral-Large-2402 1157.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253
2024-03-27 Claude 3 Opus 1253.0Elo LMSYS benchmark ranks Claude 3 Opus top with Elo rating of 1253