Researchers present VideoX-Qwen, an integrated framework combining scalable data construction with model training to advance general-purpose, instruction-driven video editing. The system utilizes a production pipeline that organizes specialized models to generate directional editing records, resulting in over 1.2 million high-quality samples across addition, removal, replacement, and attribute tasks.
- The pipeline produces more than 1.2 million directional video-editing records with an automatic acceptance rate of 89%.
- A unified Qwen-Wan editor combines multimodal semantic conditioning with dense source-video latent guidance for precise transformations.
- In comparisons with UniVideo and Kling O1, VideoX-Qwen achieved the best mean result on nine of eleven reported metrics, including instruction following and content preservation.
This work provides a practical foundation for more capable video editing by addressing the need for large-scale paired supervision and effective adaptation of video-generation backbones.