MiniMax has made its MiniMax-H3 model available on Hugging Face. This general-purpose, omni-modal generative system supports the unified understanding of multimodal contexts composed of text, images, video, and audio.

  • It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.
  • The model possesses broad multimodal context understanding and generation capabilities at the pre-training stage.
  • This design enables outstanding performance in following complex multimodal instructions.