| 2026-09-10 |
DeepSeek-V4.1-Flash |
74.2% |
DeepSeek releases V4.1-Flash model with native multimodal support
|
| 2026-09-04 |
GPT-6 Astra |
74.1% |
OpenAI releases GPT-6 Astra, a computer-use model with 1.05M context
|
| 2026-09-04 |
Muse Spark 1.3 |
75.4% |
Meta releases Muse Spark 1.3 with fewer tool calls and tokens
|
| 2026-08-26 |
GLM-5.3-Flash |
63.0% |
Ox Alpha is GLM-5.3-Flash
|
| 2026-08-22 |
GLM-5.3 |
69.0% |
GLM-5.3 matches Claude Fable 5 accuracy on DeepSWE at a fifth of the cost
|
| 2026-08-22 |
Claude Fable 5 |
69.7% |
GLM-5.3 matches Claude Fable 5 accuracy on DeepSWE at a fifth of the cost
|
| 2026-08-21 |
DeepSeek-V4-Flash-Vision-Exp |
59.3% |
DeepSeek releases DeepSeek-V4-Flash-Vision-Exp multimodal model
|
| 2026-08-20 |
Ornith-1.5 |
56.0% |
Z.ai proposes post-training scaling law and releases GLM 5.3; Ornith-1.5 launches with self-improvement
|
| 2026-08-19 |
Ornith-1.5 |
56.0% |
Ornith-1.5 releases 9B, 35B-A3B, and 397B models with self-improving training
|
| 2026-08-18 |
DeepSeek V4 Pro 0813 |
62.8% |
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
|
| 2026-08-18 |
Claude Fable 5 |
69.7% |
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
|
| 2026-08-18 |
GPT-5.6 Sol |
85.8% |
DeepSeek V4 Pro and GPT-5.6 Sol cascade solves 83% of DeepSWE tasks at $3.35
|
| 2026-08-18 |
DeepSeek V4 Pro 0813 |
88.5% |
DeepSeek V4 Pro and GPT-5.6 Sol cascade solves 83% of DeepSWE tasks at $3.35
|
| 2026-08-17 |
GLM-5.3 |
66.9% |
Z.ai launches GLM-5.3 coding model; Cursor joins SpaceXAI
|
| 2026-08-13 |
Gemini 3.7 Flash |
65.3% |
Google introduces Gemini 3.7 Flash with improved coding and lower cost
|
| 2026-08-13 |
DeepSeek-V4-Pro |
62.7% |
DeepSeek-V4-Pro GA release enhances agent capabilities and adds Responses API support
|
| 2026-08-13 |
Grok 4.6 |
65.9% |
SpaceXAI releases Grok 4.6 with 500K context and improved agentic reasoning
|
| 2026-08-07 |
DeepSeek-V4 Flash 0731 |
53.3% |
DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
|
| 2026-08-07 |
GPT-5.6 Luna |
67.2% |
DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
|
| 2026-08-03 |
Qwen3.8-Max |
56.6% |
Alibaba releases Qwen3.8-Max, a 2.4T-parameter MoE model
|
| 2026-07-31 |
DeepSeek-V4-Flash-0731 |
54.4% |
DeepSeek releases V4-Flash-0731 with major agentic gains at lower cost
|
| 2026-07-31 |
DeepSeek-V4-Flash |
54.4% |
DeepSeek releases DeepSeek-V4-Flash in public beta with enhanced agent capabilities
|
| 2026-07-27 |
Gemini 3.6 Flash |
49.0% |
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
|
| 2026-07-27 |
Kimi K3 |
68.5% |
Kimi K3 beats GPT-5.6 Sol on DeepSWE pass@4 and costs 64% less
|
| 2026-07-27 |
GPT-5.6 Sol |
72.7% |
Kimi K3 beats GPT-5.6 Sol on DeepSWE pass@4 and costs 64% less
|