벤치마크 · agentic

CursorBench

3 결과 3 모델

Cursor's private, version-tagged internal agent eval built from real Cursor sessions; the trailing number is a version, not a score. Not one of the public suites it is quoted beside.

0 19.5 39 58.5 78 2026-07-31 2026-08-16 2026-09-01 Claude Opus 5 · 70 · 2026-07-31 Grok 4.6 · 69.9 · 2026-08-13 Claude Fable 5.1 · 73.4 · 2026-09-01
Claude Opus 5 Grok 4.6 Claude Fable 5.1
타임라인
날짜 모델 점수 출처
2026-09-01 Claude Fable 5.1 73.4% Anthropic, Terminal-Bench-Science에서 52.6% 점수 및 저렴한 캐시 읽기와 함께 Claude Fable 5.1 출시
2026-08-13 Grok 4.6 69.9% SpaceXAI, 50만 토큰 컨텍스트와 개선된 에이전트 추론 기능을 갖춘 Grok 4.6 출시
2026-07-31 Claude Opus 5 70.0% Anthropic, Claude Opus 5 출시; OpenAI 모델이 보안 테스트 중 Hugging Face 침투