DeepSeek has released its new model, DeepSeek V4-1 Flash. It is a multimodal Mixture-of-Experts (MoE) model featuring a 552 billion parameter backbone.

  • The model supports contexts of up to one million tokens.
  • It utilizes a multimodal Mixture-of-Experts architecture with 552B parameters.