벤치마크 · agentic

DeepSWE

25 결과 20 모델

DataCurve's long-horizon real-repo coding-agent suite (headline metric = pass@1 resolve rate); Artificial Analysis swapped SWE-bench Pro out of its Coding Agent Index for it. Not a SWE-bench variant.

0 23.5 47 70.5 94 2026-07-27 2026-08-18 2026-09-10 Kimi K3 · 68.5 · 2026-07-27 GPT-5.6 Sol · 72.7 · 2026-07-27 GPT-5.6 Sol · 85.8 · 2026-08-18 Gemini 3.6 Flash · 49 · 2026-07-27 DeepSeek-V4-Flash · 54.4 · 2026-07-31 DeepSeek-V4-Flash-0731 · 54.4 · 2026-07-31 Qwen3.8-Max · 56.6 · 2026-08-03 DeepSeek-V4 Flash 0731 · 53.3 · 2026-08-07 GPT-5.6 Luna · 67.2 · 2026-08-07 Grok 4.6 · 65.9 · 2026-08-13 DeepSeek-V4-Pro · 62.7 · 2026-08-13 Gemini 3.7 Flash · 65.3 · 2026-08-13 GLM-5.3 · 66.9 · 2026-08-17 GLM-5.3 · 69 · 2026-08-22 DeepSeek V4 Pro 0813 · 88.5 · 2026-08-18 DeepSeek V4 Pro 0813 · 62.8 · 2026-08-18 Claude Fable 5 · 69.7 · 2026-08-18 Claude Fable 5 · 69.7 · 2026-08-22 Ornith-1.5 · 56 · 2026-08-19 Ornith-1.5 · 56 · 2026-08-20 DeepSeek-V4-Flash-Vision-Exp · 59.3 · 2026-08-21 GLM-5.3-Flash · 63 · 2026-08-26 Muse Spark 1.3 · 75.4 · 2026-09-04 GPT-6 Astra · 74.1 · 2026-09-04 DeepSeek-V4.1-Flash · 74.2 · 2026-09-10
Kimi K3 GPT-5.6 Sol Gemini 3.6 Flash DeepSeek-V4-Flash DeepSeek-V4-Flash-0731 Qwen3.8-Max DeepSeek-V4 Flash 0731 GPT-5.6 Luna Grok 4.6 DeepSeek-V4-Pro Gemini 3.7 Flash GLM-5.3 DeepSeek V4 Pro 0813 Claude Fable 5 Ornith-1.5 DeepSeek-V4-Flash-Vision-Exp GLM-5.3-Flash Muse Spark 1.3 GPT-6 Astra DeepSeek-V4.1-Flash
타임라인
날짜 모델 점수 출처
2026-09-10 DeepSeek-V4.1-Flash 74.2% DeepSeek, 네이티브 멀티모달 지원을 갖춘 V4.1-Flash 모델 출시
2026-09-04 GPT-6 Astra 74.1% OpenAI가 105만 토큰 컨텍스트를 갖춘 컴퓨터 사용 모델 GPT-6 Astra 출시
2026-09-04 Muse Spark 1.3 75.4% Meta, 도구 호출 및 토큰을 줄인 Muse Spark 1.3 출시
2026-08-26 GLM-5.3-Flash 63.0% Ox Alpha는 GLM-5.3-Flash입니다
2026-08-22 GLM-5.3 69.0% GLM-5.3은 DeepSWE에서 Claude Fable 5의 정확도를 비용의 5분의 1로 달성
2026-08-22 Claude Fable 5 69.7% GLM-5.3은 DeepSWE에서 Claude Fable 5의 정확도를 비용의 5분의 1로 달성
2026-08-21 DeepSeek-V4-Flash-Vision-Exp 59.3% DeepSeek, DeepSeek-V4-Flash-Vision-Exp 멀티모달 모델 출시
2026-08-20 Ornith-1.5 56.0% Z.ai가 사후 학습 확장 법칙을 제안하고 GLM 5.3을 출시; Ornith-1.5가 자기 개선 기능으로 데뷔
2026-08-19 Ornith-1.5 56.0% Ornith-1.5가 자기 개선 훈련을 통해 9B, 35B-A3B, 397B 모델 출시
2026-08-18 DeepSeek V4 Pro 0813 62.8% DeepSWE에서 DeepSeek V4 Pro 0813 대 Claude Fable 5: 비용, 코딩, 라우팅
2026-08-18 Claude Fable 5 69.7% DeepSWE에서 DeepSeek V4 Pro 0813 대 Claude Fable 5: 비용, 코딩, 라우팅
2026-08-18 GPT-5.6 Sol 85.8% DeepSeek V4 Pro와 GPT-5.6 Sol의 연쇄 전략이 DeepSWE 작업의 83%를 $3.35에 해결
2026-08-18 DeepSeek V4 Pro 0813 88.5% DeepSeek V4 Pro와 GPT-5.6 Sol의 연쇄 전략이 DeepSWE 작업의 83%를 $3.35에 해결
2026-08-17 GLM-5.3 66.9% Z.ai가 GLM-5.3 코딩 모델 출시; Cursor가 SpaceXAI 합류
2026-08-13 Gemini 3.7 Flash 65.3% 구글, 코딩 성능 향상 및 비용 절감을 위한 Gemini 3.7 Flash 출시
2026-08-13 DeepSeek-V4-Pro 62.7% DeepSeek-V4-Pro GA 릴리스로 에이전트 기능 강화 및 Responses API 지원 추가
2026-08-13 Grok 4.6 65.9% SpaceXAI, 50만 토큰 컨텍스트와 개선된 에이전트 추론 기능을 갖춘 Grok 4.6 출시
2026-08-07 DeepSeek-V4 Flash 0731 53.3% DeepSWE에서 DeepSeek-V4 Flash 0731 대 GPT-5.6 Luna: 비용 및 코딩
2026-08-07 GPT-5.6 Luna 67.2% DeepSWE에서 DeepSeek-V4 Flash 0731 대 GPT-5.6 Luna: 비용 및 코딩
2026-08-03 Qwen3.8-Max 56.6% 알리바바, 2.4조 파라미터 MoE 모델인 Qwen3.8-Max 출시
2026-07-31 DeepSeek-V4-Flash-0731 54.4% DeepSeek, 낮은 비용으로 주요 에이전트 성능 향상된 V4-Flash-0731 출시
2026-07-31 DeepSeek-V4-Flash 54.4% DeepSeek, 에이전트 기능 강화된 DeepSeek-V4-Flash 공개 베타 출시
2026-07-27 Gemini 3.6 Flash 49.0% 구글, Gemini 3.6 Flash, 3.5 Flash-Lite 및 3.5 Flash Cyber 출시
2026-07-27 Kimi K3 68.5% Kimi K3가 DeepSWE pass@4에서 GPT-5.6 Sol을 제치고 비용을 64% 절감
2026-07-27 GPT-5.6 Sol 72.7% Kimi K3가 DeepSWE pass@4에서 GPT-5.6 Sol을 제치고 비용을 64% 절감