LTX has released LTX-2.5, an open weights world model for video generation, real-time applications, and physical AI, optimized for local inference on NVIDIA RTX GPUs and DGX Spark hardware. The release aims to shift video production from cloud infrastructure to local desktops by reducing VRAM requirements and eliminating per-generation fees.
- Native multishot generation renders sequences as one coherent output, maintaining character and scene consistency across cuts.
- A new diffusion video decoder reduces visual artifacts in high-motion scenes while preserving compression ratios.
- On-prem generation of a 10-second clip takes 6.8 seconds on 2x NVIDIA GB200, significantly faster than closed alternatives.
- The model includes a physical AI checkpoint for robotics fine-tuning and a distilled version for lower-cost deployment.
- Weights are available on Hugging Face and ComfyUI, free for organizations under $10M in annual recurring revenue.
Local generation allows creators to batch-generate content overnight and test variations without cloud costs or IP leaving the machine.