| 2026-09-09 |
Qwen Test agent + Repair agent (post-trained) |
72.6% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
Qwen Test agent (post-trained) |
62.2% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
GPT-5.6-sol (Test agent) |
65.3% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
base Test agent (no-test baseline) |
57.3% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
base Repair agent (no-test baseline) |
61.2% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
Qwen Test agent (post-trained) + Repair agent (post-trained) |
72.6% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
Qwen Test agent (post-trained) |
62.2% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
GPT-5.6-sol (Test agent) |
65.3% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
base Repair agent (no-test baseline) |
61.2% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-09 |
base Test agent |
57.3% |
ExecCritic uses role-specific RL to improve coding agents via test-guided repair
|
| 2026-09-07 |
K2-Horizon-7B |
70.6% |
IFM releases K2 Horizon: six Apache 2.0 models from 0.9B to 375B
|
| 2026-09-07 |
K2-Horizon-3.7B |
68.6% |
IFM releases K2 Horizon: six Apache 2.0 models from 0.9B to 375B
|
| 2026-09-02 |
Qwen3.5-9B |
16.6% |
CANOPY enables Qwen3-14B to top AppWorld using outcome-only RL
|
| 2026-08-26 |
Granite 4.2 30B |
57.0% |
IBM releases Granite 4.2 with native reasoning and agentic RL for open enterprise models
|
| 2026-08-26 |
Granite 4.2 8B |
47.67% |
IBM releases Granite 4.2 with native reasoning and agentic RL for open enterprise models
|
| 2026-08-20 |
Qwen3.5-9B |
14.6% |
Meta releases Muse Video with native audio; Stripe acquires OpenRouter
|
| 2026-08-20 |
Ornith-1.5 |
86.0% |
Z.ai proposes post-training scaling law and releases GLM 5.3; Ornith-1.5 launches with self-improvement
|
| 2026-08-19 |
Ornith-1.5 |
86.0% |
Ornith-1.5 releases 9B, 35B-A3B, and 397B models with self-improving training
|
| 2026-08-18 |
Qwen3.5-27B |
74.8% |
Kozuchi Agent resolves 374 SWE-bench Verified instances using Qwen3.5-27B
|
| 2026-08-10 |
Qwen3.6-27B |
77.2% |
Meta releases Muse Glimmer, a 30B open-weights agentic model running on one consumer GPU
|
| 2026-08-03 |
Orchard-SWE |
69.7% |
Microsoft Research releases Orchard framework for scalable agentic AI
|
| 2026-08-03 |
Inkling-Small |
80.2% |
Thinking Machines Lab releases Inkling-Small, a 276B open weights multimodal MoE model
|
| 2026-07-24 |
Claude Opus 5 |
96.0% |
Anthropic releases Claude Opus 5 with default thinking and agentic coding gains
|
| 2026-07-19 |
DeepSeek V4 Pro |
80.6% |
Kimi K3, DeepSeek V4 Pro, and GLM-5.2 compared on benchmarks, license, and serving cost
|
| 2026-07-17 |
Devstral Small 2 |
68.0% |
Mistral Vibe for Code leads four coding agents on scaffold-to-PR task
|
| 2026-07-17 |
Devstral 2 |
72.2% |
Mistral Vibe for Code leads four coding agents on scaffold-to-PR task
|
| 2026-07-14 |
Qwen3-30B-A3B |
6.0% |
UMoE realigns MoE expert pools for improved domain-specific fine-tuning
|
| 2026-07-14 |
Qwen3.5-35B-A3B |
6.0% |
UMoE realigns MoE expert pools for improved domain-specific fine-tuning
|
| 2026-07-07 |
LLM-as-a-Verifier |
78.2% |
LLM-as-a-Verifier introduces general-purpose verification framework with continuous scoring
|
| 2026-06-17 |
LoopCoder-V2 |
64.4% |
LoopCoder-V2: Two-Loop PLT Model Achieves Best Gain-Cost Trade-Off
|
| 2026-06-17 |
LoopCoder-v2 |
64.4% |
LoopCoder-v2 Achieves Optimal Two-Loop Performance
|
| 2026-04-25 |
Qwen3.6-27B |
77.2% |
Alibaba releases Qwen3.6-27B, beating larger predecessor on coding benchmarks
|
| 2026-04-24 |
V4-Pro-Max |
80.6% |
DeepSeek releases V4 with efficient long-context architecture for agents
|
| 2026-03-25 |
o3 |
71.7% |
OpenAI announces o3 retirement from ChatGPT and shares benchmark scores
|
| 2026-02-17 |
Claude Sonnet 4.6 |
79.6% |
Anthropic releases Claude Sonnet 4.6 with improved coding, computer use, and 1M token context
|
| 2026-02-05 |
Claude Opus 4.6 |
80.8% |
Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
|
| 2026-02-01 |
MiniMax M2.5 |
80.2% |
Decontaminated benchmarks reveal 12-point gap between top AI coding models
|
| 2026-02-01 |
Claude Opus 4.6 |
80.8% |
Decontaminated benchmarks reveal 12-point gap between top AI coding models
|
| 2026-01-28 |
Qwen3.6-27B |
77.2% |
Qwen 3.6-27B and DeepSeek V4 Flash lead local coding benchmarks
|
| 2025-12-11 |
GPT-5.2 Thinking |
80.0% |
OpenAI releases GPT-5.2, outperforming Gemini 3 in coding and reasoning benchmarks
|
| 2025-12-11 |
GPT-5.2 Thinking |
80.0% |
OpenAI introduces GPT-5.2 with state-of-the-art professional knowledge work capabilities
|
| 2025-12-09 |
Devstral Small 2 |
68.0% |
Mistral releases Devstral 2 coding models and Mistral Vibe CLI
|
| 2025-12-09 |
Devstral 2 |
72.2% |
Mistral releases Devstral 2 coding models and Mistral Vibe CLI
|
| 2025-11-19 |
Kimi K2 Thinking |
71.3% |
Moonshot AI releases Kimi K2 Thinking with agentic tool use
|
| 2025-11-19 |
GPT-5.1-Codex-Max |
77.9% |
OpenAI makes GPT-5.1-Codex-Max the default in Codex CLI
|
| 2025-11-07 |
Kimi K2 Thinking |
71.3% |
Moonshot AI releases Kimi K2 Thinking, an open-source 1T parameter reasoning model
|
| 2025-09-30 |
Claude Sonnet 4.5 |
77.2% |
Anthropic releases Claude Sonnet 4.5 with improved coding and agent capabilities
|
| 2025-09-29 |
Claude Sonnet 4.5 |
77.2% |
Anthropic
|
| 2025-08-07 |
GPT-5 |
74.9% |
OpenAI launches unified GPT-5 with real-time routing and free access
|
| 2025-08-07 |
GPT-5 |
74.9% |
OpenAI introduces GPT-5 with unified routing and expert-level reasoning
|
| 2025-08-07 |
GPT‑5 |
74.9% |
OpenAI releases GPT-5 API with coding and agentic improvements
|
| 2025-08-07 |
GPT-5 |
74.9% |
OpenAI launches GPT-5 with adaptive reasoning and unified architecture
|
| 2025-08-07 |
GPT-5 |
74.9% |
OpenAI
|
| 2025-08-06 |
GLM-4.5 |
64.2% |
Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek
|
| 2025-07-28 |
Kimi K2 |
65.8% |
Kimi K2: open MoE model with 1T parameters and MuonClip optimizer
|
| 2025-07-28 |
Kimi K2 |
65.8% |
Moonshot releases Kimi K2, an open agentic MoE model with 32B activated parameters
|
| 2025-07-11 |
Kimi K2-Instruct |
65.8% |
Moonshot AI releases Kimi-K2-Instruct, a 1T-parameter MoE model with agentic capabilities
|
| 2025-07-11 |
Kimi K2 |
65.8% |
Moonshot AI releases Kimi K2, a 1T-parameter MoE model with agentic capabilities
|
| 2025-07-10 |
Devstral Small 1.1 |
53.6% |
Mistral AI releases Devstral Small 1.1 and Devstral Medium with All Hands AI
|
| 2025-07-10 |
Devstral Medium |
61.6% |
Mistral AI releases Devstral Small 1.1 and Devstral Medium with All Hands AI
|
| 2025-05-30 |
Deepseek-R1-0528 |
57.6% |
DeepSeek updates R1 model with improved reasoning and benchmark scores
|
| 2025-05-23 |
Claude Opus 4 |
72.5% |
Anthropic releases Claude Opus 4 and Sonnet 4 with hybrid reasoning
|
| 2025-05-23 |
Claude Sonnet 4 |
72.7% |
Anthropic releases Claude Opus 4 and Sonnet 4 with hybrid reasoning
|
| 2025-05-22 |
Claude Opus 4 |
72.5% |
Anthropic
|
| 2025-05-21 |
Devstral |
46.8% |
Mistral AI releases Devstral agentic LLM for software engineering
|
| 2025-04-17 |
o4-mini |
68.1% |
OpenAI releases o4-mini, a faster, cheaper multimodal reasoning model
|
| 2025-04-16 |
OpenAI o3 |
71.7% |
OpenAI
|
| 2025-04-14 |
GPT‑4.1 |
54.6% |
OpenAI launches GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano in the API
|
| 2025-04-14 |
GPT-4.1 |
54.6% |
OpenAI releases GPT-4.1 family with 1M token context and coding improvements
|
| 2025-03-26 |
Gemini 2.5 Pro |
63.8% |
Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
|
| 2025-03-25 |
Gemini 2.5 Pro |
63.8% |
Google introduces Gemini 2.5 Pro thinking model
|
| 2025-03-25 |
Gemini 2.5 Pro |
63.8% |
Google introduces Gemini 2.5 Pro experimental model
|
| 2025-03-25 |
Gemini 2.5 Pro |
63.8% |
Google DeepMind
|
| 2025-03-05 |
GPT-4.5 |
38.0% |
OpenAI releases GPT-4.5, its largest model, for ChatGPT Plus
|
| 2025-02-24 |
Claude 3.7 Sonnet |
62.3% |
Anthropic
|
| 2025-01-06 |
Claude 3.5 Sonnet |
49.0% |
Claude 3.5 Sonnet achieves 49% on SWE-bench Verified
|
| 2024-12-20 |
o3 |
71.7% |
OpenAI previews o3 reasoning model with ARC AGI and Frontier Math breakthroughs
|
| 2024-12-20 |
o3 |
71.7% |
OpenAI releases o3 model with high performance in reasoning and coding
|
| 2024-10-30 |
Claude 3.5 Sonnet |
49.0% |
Claude 3.5 Sonnet achieves 49% on SWE-bench Verified with minimal agent scaffolding
|
| 2024-10-22 |
Claude 3.5 Haiku |
40.6% |
Anthropic upgrades Claude 3.5 Sonnet, releases Claude 3.5 Haiku and computer use beta
|
| 2024-10-22 |
Claude 3.5 Sonnet |
49.0% |
Anthropic upgrades Claude 3.5 Sonnet, releases Claude 3.5 Haiku and computer use beta
|
| 2024-10-22 |
Claude 3.5 Sonnet |
49.0% |
Anthropic
|
| 2024-08-13 |
Agentless |
32.0% |
OpenAI releases SWE-bench Verified with human-validated subset
|
| 2024-08-13 |
GPT‑4o |
33.2% |
OpenAI releases SWE-bench Verified with human-validated subset
|