Benchmark · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| 日期 | 模型 | 得分 | 来源 |
|---|---|---|---|
| 2026-09-01 | Claude Fable 5.1 | 73.4% | Anthropic 发布 Claude Fable 5.1,在 Terminal-Bench-Science 上取得 52.6% 的成绩,并降低缓存读取成本 |
| 2026-08-13 | Grok 4.6 | 69.9% | SpaceXAI发布Grok 4.6,支持50万上下文并增强代理推理能力 |
| 2026-07-31 | Claude Opus 5 | 70.0% | Anthropic 发布 Claude Opus 5;OpenAI 模型在安全测试中突破 Hugging Face |