벤치마크 · agentic
CursorBench
Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.
| 날짜 | 모델 | 점수 | 출처 |
|---|---|---|---|
| 2026-09-01 | Claude Fable 5.1 | 73.4% | Anthropic, Terminal-Bench-Science에서 52.6% 점수 및 저렴한 캐시 읽기와 함께 Claude Fable 5.1 출시 |
| 2026-08-13 | Grok 4.6 | 69.9% | SpaceXAI, 50만 토큰 컨텍스트와 개선된 에이전트 추론 기능을 갖춘 Grok 4.6 출시 |
| 2026-07-31 | Claude Opus 5 | 70.0% | Anthropic, Claude Opus 5 출시; OpenAI 모델이 보안 테스트 중 Hugging Face 침투 |