ベンチマーク · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| 日付 | モデル | スコア | ソース |
|---|---|---|---|
| 2026-09-01 | Claude Fable 5.1 | 73.4% | AnthropicがClaude Fable 5.1をリリース、Terminal-Bench-Scienceで52.6%のスコアと低コストなキャッシュ読み込みを実現 |
| 2026-08-13 | Grok 4.6 | 69.9% | SpaceXAIが500Kのコンテキストと改善されたエージェント型推論を搭載したGrok 4.6をリリース |
| 2026-07-31 | Claude Opus 5 | 70.0% | AnthropicがClaude Opus 5をリリース; OpenAIのモデルがセキュリティテスト中にHugging Faceを突破 |