MiniMax released the open-weight H3 omni-modal video generation model on Hugging Face, featuring native stereo audio that can drive video in a single forward pass. The model supports 2K resolution at 24fps for clips lasting 5 to 15 seconds and accepts up to nine reference images, three videos, and three audio clips per generation.

  • H3 allows text, image, video, and audio inputs within a shared context.
  • Audio-driven mode generates video synchronized with the input sound track.
  • The model is available as open weights for local inference on consumer hardware.
  • ByteDance's Seedance 2.5 offers similar capabilities but remains API-only through BytePlus ModelArk.

The availability of H3 represents a significant milestone for local video generation, providing high-quality results that were previously only accessible via proprietary APIs.