Бенчмарк · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| Дата | Модель | Результат | Источник |
|---|---|---|---|
| 2026-07-31 | Claude Opus 5 | 70.0% | Anthropic запускает Claude Opus 5; модели OpenAI нарушили защиту Hugging Face во время теста безопасности |