| 2026-09-10 |
DeepSeek-V4.1-Flash |
74.2% |
DeepSeek发布具有原生多模态支持的V4.1-Flash模型
|
| 2026-09-04 |
GPT-6 Astra |
74.1% |
OpenAI发布GPT-6 Astra,一款拥有1.05M上下文的计算机使用模型
|
| 2026-09-04 |
Muse Spark 1.3 |
75.4% |
Meta 发布 Muse Spark 1.3,减少工具调用和 token 数量
|
| 2026-08-26 |
GLM-5.3-Flash |
63.0% |
Ox Alpha 即 GLM-5.3-Flash
|
| 2026-08-22 |
GLM-5.3 |
69.0% |
GLM-5.3 在 DeepSWE 上以五分之一的成本达到与 Claude Fable 5 相当的准确率
|
| 2026-08-22 |
Claude Fable 5 |
69.7% |
GLM-5.3 在 DeepSWE 上以五分之一的成本达到与 Claude Fable 5 相当的准确率
|
| 2026-08-21 |
DeepSeek-V4-Flash-Vision-Exp |
59.3% |
DeepSeek 发布 DeepSeek-V4-Flash-Vision-Exp 多模态模型
|
| 2026-08-20 |
Ornith-1.5 |
56.0% |
Z.ai提出后训练缩放定律并发布GLM 5.3;Ornith-1.5以自我改进能力上线
|
| 2026-08-19 |
Ornith-1.5 |
56.0% |
Ornith-1.5 发布具有自我改进训练功能的 9B、35B-A3B 和 397B 模型
|
| 2026-08-18 |
DeepSeek V4 Pro 0813 |
62.8% |
DeepSeek V4 Pro 0813 与 Claude Fable 5 在 DeepSWE 上的对比:成本、编码与路由
|
| 2026-08-18 |
Claude Fable 5 |
69.7% |
DeepSeek V4 Pro 0813 与 Claude Fable 5 在 DeepSWE 上的对比:成本、编码与路由
|
| 2026-08-18 |
GPT-5.6 Sol |
85.8% |
DeepSeek V4 Pro 与 GPT-5.6 Sol 级联策略在 DeepSWE 任务上实现 83% 通过率,成本为 3.35 美元
|
| 2026-08-18 |
DeepSeek V4 Pro 0813 |
88.5% |
DeepSeek V4 Pro 与 GPT-5.6 Sol 级联策略在 DeepSWE 任务上实现 83% 通过率,成本为 3.35 美元
|
| 2026-08-17 |
GLM-5.3 |
66.9% |
Z.ai 发布 GLM-5.3 编码模型;Cursor 加入 SpaceXAI
|
| 2026-08-13 |
Gemini 3.7 Flash |
65.3% |
Google推出Gemini 3.7 Flash,提升编码能力并降低成本
|
| 2026-08-13 |
DeepSeek-V4-Pro |
62.7% |
DeepSeek-V4-Pro GA 发布增强代理能力并添加 Responses API 支持
|
| 2026-08-13 |
Grok 4.6 |
65.9% |
SpaceXAI发布Grok 4.6,支持50万上下文并增强代理推理能力
|
| 2026-08-07 |
DeepSeek-V4 Flash 0731 |
53.3% |
DeepSeek-V4 Flash 0731 与 GPT-5.6 Luna 在 DeepSWE 上的对比:成本与编码
|
| 2026-08-07 |
GPT-5.6 Luna |
67.2% |
DeepSeek-V4 Flash 0731 与 GPT-5.6 Luna 在 DeepSWE 上的对比:成本与编码
|
| 2026-08-03 |
Qwen3.8-Max |
56.6% |
阿里巴巴发布Qwen3.8-Max,一款拥有2.4万亿参数的MoE模型
|
| 2026-07-31 |
DeepSeek-V4-Flash-0731 |
54.4% |
DeepSeek 发布 V4-Flash-0731,以更低的成本实现显著的智能体能力提升
|
| 2026-07-31 |
DeepSeek-V4-Flash |
54.4% |
DeepSeek 发布 DeepSeek-V4-Flash 公测版,增强智能体能力
|
| 2026-07-27 |
Gemini 3.6 Flash |
49.0% |
Google发布Gemini 3.6 Flash、3.5 Flash-Lite和3.5 Flash Cyber
|
| 2026-07-27 |
Kimi K3 |
68.5% |
Kimi K3在DeepSWE pass@4上超越GPT-5.6 Sol,成本降低64%
|
| 2026-07-27 |
GPT-5.6 Sol |
72.7% |
Kimi K3在DeepSWE pass@4上超越GPT-5.6 Sol,成本降低64%
|