| 2026-09-10 |
DeepSeek-V4.1-Flash |
90.6% |
DeepSeek releases V4.1-Flash model with native multimodal support
|
| 2026-09-07 |
K2-Horizon-MoVA-36B-A4B |
58.6% |
IFM releases K2 Horizon: six Apache 2.0 models from 0.9B to 375B
|
| 2026-09-07 |
K2-Horizon-375B-A23B |
66.9% |
IFM releases K2 Horizon: six Apache 2.0 models from 0.9B to 375B
|
| 2026-09-04 |
Muse Spark 1.3 |
88.8% |
Meta releases Muse Spark 1.3 with fewer tool calls and tokens
|
| 2026-08-26 |
Granite 4.2 30B |
29.24% |
IBM releases Granite 4.2 with native reasoning and agentic RL for open enterprise models
|
| 2026-08-26 |
Granite 4.2 8B |
20.56% |
IBM releases Granite 4.2 with native reasoning and agentic RL for open enterprise models
|
| 2026-08-21 |
Grok 4.6 |
88.4% |
SpaceXAI introduces Grok 4.6 with Cursor data for agentic work
|
| 2026-08-21 |
DeepSeek-V4-Flash-Vision-Exp |
83.9% |
DeepSeek releases DeepSeek-V4-Flash-Vision-Exp multimodal model
|
| 2026-08-20 |
agents |
94.0% |
Meta releases Muse Video with native audio; Stripe acquires OpenRouter
|
| 2026-08-19 |
Ornith-1.5 |
86.1% |
Ornith-1.5 releases 9B, 35B-A3B, and 397B models with self-improving training
|
| 2026-08-13 |
DeepSeek-V4-Pro |
87.9% |
DeepSeek-V4-Pro GA release enhances agent capabilities and adds Responses API support
|
| 2026-08-13 |
Grok 4.6 |
88.4% |
xAI releases Grok 4.6; Alibaba opens Qwen3.8-MoE; DeepSeek V4 Pro GA
|
| 2026-08-11 |
Opus 5 |
86.74% |
Ouroboros self-developing agent sets new benchmarks with Opus 5
|
| 2026-08-10 |
Qwen3.6-27B |
60.7% |
Meta releases Muse Glimmer, a 30B open-weights agentic model running on one consumer GPU
|
| 2026-08-09 |
DeepSeek V4 Flash 0731 |
82.7% |
DeepSeek V4 Flash 0731 matches reported 82.7% on Terminal-Bench 2.1 in public harness
|
| 2026-08-08 |
Pokee-Isaac 28B |
65.1% |
Pokee AI releases Pokee-Isaac 28B, a 10M-token context model for in-boundary deployment
|
| 2026-08-07 |
DeepSeek-V4-Flash-0731 |
82.7% |
DeepSeek releases V4-Flash-0731, outperforming V4-Pro at lower cost
|
| 2026-08-05 |
Muse Spark 1.1 |
80.0% |
Meta releases Muse Code beta terminal agent powered by Muse Spark 1.2
|
| 2026-08-05 |
Qwen3.8-Max |
67.4% |
Alibaba announces Qwen3.8-Max with 2.4T parameters and upcoming open weights
|
| 2026-08-03 |
Qwen3.8-Max |
86.6% |
Alibaba releases Qwen3.8-Max, a 2.4T-parameter MoE model
|
| 2026-08-03 |
Inkling-Small |
64.7% |
Thinking Machines Lab releases Inkling-Small, a 276B open weights multimodal MoE model
|
| 2026-08-01 |
V4 Flash 0731 |
79.0% |
DeepSeek releases V4-Flash with major post-training gains and open weights
|
| 2026-08-01 |
DeepSeek-V4-Flash |
79.0% |
DeepSeek releases V4-Flash with post-training gains and open weights
|
| 2026-07-31 |
DeepSeek-V4-Flash-0731 |
82.7% |
DeepSeek releases V4-Flash-0731 with major agentic gains at lower cost
|
| 2026-07-31 |
DeepSeek-V4-Flash |
82.7% |
DeepSeek releases DeepSeek-V4-Flash in public beta with enhanced agent capabilities
|
| 2026-07-27 |
Gemini 3.5 Flash-Lite |
54.0% |
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
|
| 2026-07-26 |
KAT-Coder-V2.5 |
60.7% |
KwaiKAT releases KAT-Coder-V2.5, an agentic coding model trained on 100,000+ verifiable environments
|
| 2026-07-21 |
Gemini 3.5 Flash-Lite |
54.0% |
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for agentic workloads
|
| 2026-07-21 |
Gemini 3.5 Flash-Lite |
54.0% |
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
|
| 2026-07-08 |
hermes |
55.0% |
User benchmarks Mimo v2.5 against DeepSeek V4 Flash and Hermes
|
| 2026-07-08 |
Grok 4.5 |
83.3% |
xAI releases Grok 4.5 with low cost and mixed benchmark performance
|
| 2026-07-07 |
LLM-as-a-Verifier |
86.5% |
LLM-as-a-Verifier introduces general-purpose verification framework with continuous scoring
|
| 2026-07-05 |
GLM 5.2 FP8 with FP8 KV |
79.8% |
GLM 5.2 FP8 with FP8 KV achieves 79.8% on Terminal-Bench 2.1
|
| 2026-06-23 |
Tmax-27B |
43.0% |
Tmax-27B Terminal Agent for Small GPUs with DPPO Training
|
| 2026-06-23 |
Tmax |
27.0% |
Tmax: A Simple RL Recipe for Terminal Agents
|
| 2026-04-25 |
Qwen3.6-27B |
59.3% |
Alibaba releases Qwen3.6-27B, beating larger predecessor on coding benchmarks
|
| 2026-04-25 |
GPT-5.5 |
82.7% |
OpenAI releases GPT-5.5, a ground-up retrained model with native omnimodal processing
|
| 2026-04-23 |
GPT-5.5 |
82.7% |
OpenAI releases GPT-5.5 with agentic capabilities and doubled API pricing
|
| 2026-04-23 |
GPT-5.5 |
82.7% |
OpenAI releases GPT-5.5 and GPT-5.5 Pro in API
|
| 2026-04-12 |
MiniMax M2.7 |
57.0% |
MiniMax open-sources M2.7, a self-evolving MoE model scoring 56.22% on SWE-Pro
|
| 2026-02-10 |
Step 3.5 Flash |
51.0% |
Step 3.5 Flash: open 11B-active MoE model achieves frontier-level agentic performance
|
| 2025-11-19 |
GPT-5.1-Codex-Max |
58.1% |
OpenAI makes GPT-5.1-Codex-Max the default in Codex CLI
|