Benchmark · agentic
Terminal-Bench 2
Harbor/Terminus-2 89-task terminal-agent suite; 2.0 and 2.1 are the same task set, so one series. File here only when the quote says 2.0 / 2.1 / v2 — not the original suite (terminal-bench), and not the harder 'Terminal-Bench Hard' subset (not in this catalog).
| 日期 | 模型 | 得分 | 来源 |
|---|---|---|---|
| 2026-07-21 | Gemini 3.5 Flash-Lite | 54.0% | Google发布Gemini 3.6 Flash、3.5 Flash-Lite和3.5 Flash Cyber,面向智能体工作负载 |
| 2026-07-21 | Gemini 3.5 Flash-Lite | 54.0% | Google发布Gemini 3.6 Flash、3.5 Flash-Lite和3.5 Flash Cyber |
| 2026-07-08 | hermes | 55.0% | 用户将 Mimo v2.5 与 DeepSeek V4 Flash 和 Hermes 进行基准测试对比 |
| 2026-07-08 | Grok 4.5 | 83.3% | xAI发布Grok 4.5,成本低且基准测试表现参差不齐 |
| 2026-07-07 | LLM-as-a-Verifier | 86.5% | LLM-as-a-Verifier 引入具有连续评分的通用验证框架 |
| 2026-07-05 | GLM 5.2 FP8 with FP8 KV | 79.8% | GLM 5.2 FP8 配合 FP8 KV 在 Terminal-Bench 2.1 上达到 79.8% |
| 2026-04-25 | Qwen3.6-27B | 59.3% | 阿里巴巴发布Qwen3.6-27B,在编程基准测试中超越更大的前代模型 |
| 2026-04-25 | GPT-5.5 | 82.7% | OpenAI发布GPT-5.5,一款从头重新训练、原生支持全模态处理的模型 |
| 2026-04-23 | GPT-5.5 | 82.7% | OpenAI 通过 API 发布 GPT-5.5 和 GPT-5.5 Pro |
| 2026-04-23 | GPT-5.5 | 82.7% | OpenAI发布具备智能体能力的GPT-5.5,API定价翻倍 |
| 2026-04-12 | MiniMax M2.7 | 57.0% | MiniMax开源M2.7,一款在SWE-Pro上得分56.22%的自进化MoE模型 |
| 2026-02-10 | Step 3.5 Flash | 51.0% | 步骤 3.5 Flash:开源 11B 活跃参数的 MoE 模型达到前沿级代理性能 |
| 2025-11-19 | GPT-5.1-Codex-Max | 58.1% | OpenAI 将 GPT-5.1-Codex-Max 设为 Codex CLI 的默认模型 |