Topic · Video generation
media MarkTechPost · 8d ago · 15 views

MiniMax releases MiniMax H3, an omni-modal video model generating 2K clips with native stereo audio

MiniMax has released MiniMax H3, a general-purpose multimodal generation model that unifies text, image, video, and audio into a single context to produce 2K video clips up to 15 seconds long with native stereo sound. Unlike previous stacks that rely on separate expert models for specific tasks, H3 allows users to express reference and editing relationships through natural language prompts.

lab Hugging Face Blog · 13d ago · 19 views

NVIDIA introduces Cosmos-H-Dreams, a real-time generative simulator for surgical robotics

NVIDIA has introduced Cosmos-H-Dreams, a real-time, action-conditioned generative simulator for surgical robotics that distills the capabilities of its Cosmos-H-Surgical-Simulator into a causal student model. Served through NVIDIA's FlashDreams accelerated streaming-inference library, the system enables interactive control by a person or learned policy in a closed loop.

arxiv arXiv cs.CL · 24d ago

Text2Sign: Single-GPU diffusion baseline for text-to-sign language video generation

Researchers present Text2Sign, a text-conditioned diffusion model designed to generate short sign-language clips on a single NVIDIA L4 GPU, addressing the high costs associated with training and evaluating video diffusion models. The system combines a frozen vision-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce computational requirements while preserving motion coherence.