MiniMax has launched MiniMax H3, a general-purpose multimodal generation model that understands unified context across text, images, video, and audio. The model generates video with native stereo sound at up to 15 seconds duration and 2K resolution.

  • H3 features Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-Context Regeneration technologies.
  • It excels at instruction following, accurate text and brand rendering, and V2V motion transfer.
  • At 2K resolution, the per-second price is less than a third of mainstream models; at 768p, it is less than half the price of mainstream models' 720p.
  • MiniMax plans to release open weights in the coming days to support the open-source community and accelerate hardware compatibility.

The release aims to address the dominance of closed-source video generation models by providing an open ecosystem and customizable versions for users.