DeepSeek has released DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT-4 Turbo in code-specific tasks. The model is further pre-trained from an intermediate checkpoint of DeepSeek-V2 using an additional 6 trillion tokens.

  • Enhances coding and mathematical reasoning capabilities while maintaining general language performance.
  • Expands programming language support from 86 to 338 languages.
  • Extends context length from 16K to 128K.
  • Outperforms closed-source models like GPT-4 Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks.

This release provides a high-performance open alternative for code intelligence, offering significant advancements over its predecessor DeepSeek-Coder-33B.