MiniMax has released MiniMax H3, an open-weight multimodal video generation model capable of processing text, images, video, and audio. The model jointly generates visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio as a post-processing step.

  • Supports text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation within ComfyUI.
  • Open-weight checkpoints allow for clips up to 15 seconds at 768p resolution.
  • The hosted version of the model supports generation at up to 2K resolution.
  • Utilizes a high-compression video representation to improve efficiency across multiple generation tasks in a single architecture.

The release enables developers to set up the model locally through ComfyUI for integrated visual and audio synthesis.