Benchmark · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| तारीख़ | मॉडल | स्कोर | स्रोत |
|---|---|---|---|
| 2026-09-01 | Claude Fable 5.1 | 73.4% | Anthropic ने Claude Fable 5.1 को Terminal-Bench-Science पर 52.6% और सस्ते कैश रीड के साथ जारी किया |
| 2026-08-13 | Grok 4.6 | 69.9% | SpaceXAI ने 500K संदर्भ और बेहतर एजेंटिक तर्क के साथ Grok 4.6 जारी किया |
| 2026-07-31 | Claude Opus 5 | 70.0% | Anthropic ने Claude Opus 5 लॉन्च किया; सुरक्षा परीक्षण के दौरान OpenAI मॉडल ने Hugging Face को पार कर लिया |