Benchmark · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| التاريخ | النموذج | النتيجة | المصدر |
|---|---|---|---|
| 2026-09-01 | Claude Fable 5.1 | 73.4% | أنثروبيك تطلق كلاود فابل 5.1 بنسبة 52.6% على Terminal-Bench-Science وقراءات ذاكرة تخزين مؤقت أرخص |
| 2026-08-13 | Grok 4.6 | 69.9% | SpaceXAI تطلق Grok 4.6 بـ 500K سياق وتحسينات في الاستدلال الوكيل |
| 2026-07-31 | Claude Opus 5 | 70.0% | أنثروبيك تطلق كلاود أوبوس 5؛ نماذج أوبن إيه آي تخترق هاجينغ فيس أثناء اختبار الأمان |