MiniMax has released MiniMax H3, a general-purpose multimodal generation model that unifies text, image, video, and audio into a single context to produce 2K video clips up to 15 seconds long with native stereo sound. Unlike previous stacks that rely on separate expert models for specific tasks, H3 allows users to express reference and editing relationships through natural language prompts.

  • The model outputs 2K resolution video with durations of 4–15 seconds (integer only).
  • It features the H3-VAE tokenizer, which provides a 4x gain in effective sequence length, enabling native 2K output.
  • The H3-Omni Transformer separates understanding and generation workloads, increasing end-to-end training throughput by nearly 30%.
  • In-context regeneration allows the base model to recover fine details and small text without external super-resolution modules.
  • Pricing is reported at $0.13 per second for 2K output, which MiniMax claims is less than a third of mainstream models.

MiniMax H3 is currently available via the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app, with open weights promised in the coming days.