DeepSeek has released a significant update to its DeepSeek-R1-0528 model, enhancing its reasoning capabilities and closing the performance gap with leading proprietary models from OpenAI and Google. The update utilizes improved algorithms and increased computing power to boost accuracy across math, coding, and logic benchmarks without altering the core architecture.

  • On AIME 2025, accuracy rose from 70% to 87.5%, while average prompt tokens increased from 12,000 to 23,000.
  • Math benchmarks saw gains on AIME 2024 (79.8% to 91.4%), HMMT 2025 (41.7% to 79.4%), and CNMO 2024 (78.8% to 86.9%).
  • Coding performance improved on LiveCodeBench (63.5% to 73.3%) and Aider-Polyglot (53.3% to 71.6%), with a Codeforces rating increase from 1530 to 1930.
  • General knowledge scores increased on GPQA-Diamond (71.5% to 81.0%) and Humanity's Last Exam (8.5% to 17.7%).
  • Independent platform Artificial Analysis rated the model at 68, ranking it ahead of Meta's Llama 4 Maverick and Alibaba's Qwen3 253.

The update demonstrates that open-weight models can achieve competitive performance through increased post-training with reinforcement learning. DeepSeek also released a distilled version, DeepSeek-R1-0528-Qwen3-8B, which achieves 86% on AIME 2024 while running efficiently on an Nvidia H100.