| 2026-09-11 |
GPT-6 Astra |
96.3% |
OpenAI launches GPT-6 Astra, tops leaderboards with lower cost
|
| 2026-09-10 |
DeepSeek-V4.1-Flash |
90.9% |
DeepSeek releases V4.1-Flash model with native multimodal support
|
| 2026-09-07 |
K2-Horizon-375B-A23B |
87.3% |
IFM releases K2 Horizon: six Apache 2.0 models from 0.9B to 375B
|
| 2026-08-28 |
GLM-5.3 |
91.72% |
Z.ai fine-tunes GLM-5.3 for cybersecurity gains
|
| 2026-08-25 |
QAH model (GPT-OSS 120B compressed to 60B, MXFP4) |
67.4% |
Multiverse Computing's QAH makes 4-bit GPT-OSS outperform its full-precision original
|
| 2026-08-21 |
Grok 4.6 |
94.9% |
SpaceXAI introduces Grok 4.6 with Cursor data for agentic work
|
| 2026-08-03 |
Qwen3.8-Max |
92.6% |
Alibaba releases Qwen3.8-Max, a 2.4T-parameter MoE model
|
| 2026-08-03 |
Inkling-Small |
89.5% |
Thinking Machines Lab releases Inkling-Small, a 276B open weights multimodal MoE model
|
| 2026-07-17 |
Kimi K3 |
93.5% |
Moonshot AI releases open Kimi K3, a 2.8T-parameter MoE model with 1M context
|
| 2026-07-17 |
GPT-Live-1 |
84.2% |
OpenAI releases full-duplex voice models and German court holds Google liable for AI Overviews
|
| 2026-03-25 |
o3 |
87.7% |
OpenAI announces o3 retirement from ChatGPT and shares benchmark scores
|
| 2026-02-25 |
Grok 4 |
87.7% |
xAI releases Grok 4 reasoning model with strong agentic and coding benchmarks
|
| 2026-02-05 |
Claude Opus 4.6 |
91.3% |
Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
|
| 2026-01-27 |
Kimi K2.5 |
87.6% |
Moonshot releases Kimi K2.5 with Agent Swarm technology
|
| 2025-12-11 |
GPT-5.2 Thinking |
92.4% |
OpenAI releases GPT-5.2 Pro and Thinking models for science and math
|
| 2025-12-11 |
GPT-5.2 Pro |
93.2% |
OpenAI introduces GPT-5.2 with improved coding, long-context, and reasoning capabilities
|
| 2025-12-11 |
GPT-5.2 Pro |
93.2% |
OpenAI releases GPT-5.2 Pro and Thinking models for science and math
|
| 2025-12-11 |
GPT-5.2 Thinking |
92.4% |
OpenAI introduces GPT-5.2 with improved coding, long-context, and reasoning capabilities
|
| 2025-12-04 |
Gemini 3 Pro |
93.0% |
Epoch AI launches Frontier Data Centers Hub and analyzes OSWorld benchmark
|
| 2025-11-18 |
Gemini 3 Deep Think |
93.8% |
Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
|
| 2025-11-18 |
Gemini 3 Deep Think |
93.8% |
Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
|
| 2025-11-18 |
Gemini 3 Pro |
91.9% |
Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
|
| 2025-11-18 |
Gemini 3 Pro |
91.9% |
Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
|
| 2025-10-24 |
O1 |
77.0% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
Claude Opus 4.6 |
91.3% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
GPT-5.2 |
92.4% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
Gemini 3 Pro |
91.9% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
Aristotle-X1 |
92.4% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
Gemini 3.1 Pro Preview |
94.1% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
GPT-4 |
39.0% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-10-24 |
Claude 3 Opus |
60.0% |
GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
|
| 2025-09-30 |
Claude Sonnet 4.5 |
83.4% |
Anthropic releases Claude Sonnet 4.5 with improved coding and agent capabilities
|
| 2025-09-22 |
DeepSeek-V3.1-Terminus |
74.24% |
DeepSeek releases DeepSeek-V3.1-Terminus with language and agent fixes
|
| 2025-08-07 |
GPT-5 pro |
89.4% |
OpenAI launches unified GPT-5 with real-time routing and free access
|
| 2025-08-07 |
GPT-5 pro |
88.4% |
OpenAI introduces GPT-5 with unified routing and expert-level reasoning
|
| 2025-08-07 |
GPT-5 Pro |
88.4% |
OpenAI launches GPT-5 with adaptive reasoning and unified architecture
|
| 2025-07-28 |
Kimi K2 |
75.1% |
Kimi K2: open MoE model with 1T parameters and MuonClip optimizer
|
| 2025-07-28 |
Kimi K2 |
75.1% |
Moonshot releases Kimi K2, an open agentic MoE model with 32B activated parameters
|
| 2025-05-30 |
Deepseek-R1-0528 |
81.0% |
DeepSeek updates R1 model with improved reasoning and benchmark scores
|
| 2025-05-22 |
Claude Opus 4 |
74.9% |
Anthropic releases Claude Opus 4 and Sonnet 4 with ASL-3 safety measures
|
| 2025-05-22 |
Claude Opus 4 |
74.9% |
Anthropic introduces Claude Opus 4 and Sonnet 4 with extended thinking and tool use
|
| 2025-05-22 |
Claude Sonnet 4 |
70.0% |
Anthropic introduces Claude Opus 4 and Sonnet 4 with extended thinking and tool use
|
| 2025-04-16 |
OpenAI o3 |
87.7% |
OpenAI
|
| 2025-04-15 |
Gemini 2.0 Flash |
60.1% |
Google releases Gemini 2.0 Flash with enhanced quality and twice the speed of Gemini 1.5 Pro
|
| 2025-04-14 |
GPT‑4.1 nano |
50.3% |
OpenAI launches GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano in the API
|
| 2025-04-10 |
Seed1.5-Thinking |
77.3% |
Seed1.5-Thinking uses reinforcement learning to improve reasoning
|
| 2025-04-05 |
Gemini 2.5 Pro |
84.0% |
Google expands access to Gemini 2.5 Pro amid strong benchmark results
|
| 2025-03-26 |
Gemini 2.5 Pro |
84.0% |
Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
|
| 2025-03-25 |
Gemini 2.5 Pro (experimental) |
84.0% |
Google DeepMind releases Gemini 2.5 Pro experimental model
|
| 2025-03-25 |
Gemini 2.5 Pro |
84.0% |
Google DeepMind
|
| 2025-03-24 |
DeepSeek-V3-0324 |
68.4% |
DeepSeek-V3-0324 improves reasoning, coding, and Chinese writing over DeepSeek-V3
|
| 2025-03-11 |
Grok 3 |
84.6% |
xAI releases Grok 3 with deep reasoning and million-token context
|
| 2025-03-05 |
GPT-4.5 |
71.4% |
OpenAI releases GPT-4.5, its largest model, for ChatGPT Plus
|
| 2025-02-26 |
Claude 3.5 Sonnet |
65.0% |
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
|
| 2025-02-26 |
Claude 3.7 Sonnet |
84.8% |
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
|
| 2025-02-26 |
Grok 3 Beta |
84.6% |
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
|
| 2025-02-26 |
DeepSeek R1 |
71.5% |
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
|
| 2025-02-26 |
OpenAI o3-mini |
78.0% |
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
|
| 2025-02-24 |
Claude 3.7 Sonnet |
84.8% |
Anthropic releases Claude 3.7 Sonnet with dynamic reasoning control
|
| 2025-02-19 |
Grok 3 (Think) |
84.6% |
xAI releases Grok 3 Beta and DeepSearch agent
|
| 2024-12-20 |
o3 |
87.7% |
OpenAI releases o3 model with high performance in reasoning and coding
|
| 2024-11-28 |
QwQ-32B-Preview |
65.2% |
Qwen releases QwQ-32B-Preview, an experimental reasoning model
|
| 2024-07-15 |
Qwen2-72B |
37.9% |
Qwen Team releases Qwen2 series with 72B dense and MoE models
|
| 2024-07-15 |
Qwen2-72B |
37.9% |
Qwen2 Technical Report introduces dense and MoE models up to 72B
|
| 2024-06-20 |
Claude 3.5 Sonnet |
59.4% |
Anthropic releases Claude 3.5 Sonnet addendum with improved coding and vision benchmarks
|
| 2024-06-20 |
Claude 3.5 Sonnet |
59.4% |
Anthropic
|
| 2023-11-20 |
GPT-4 based baseline |
39.0% |
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
|