Alibaba Cloud has introduced Qwen3.8-Omni-Flash, a natively multimodal agentic model designed for real-world productivity tasks that substantially improves multimodal understanding and reasoning compared to previous omni models.

The model inherits the sparse mixture-of-experts architecture from Qwen3.8-Next and extends the context window to one million tokens to support long-context multimodal reasoning and long-horizon planning. It utilizes a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities to audio and video tasks.

To support deployment, the team releases Qwen-MM-Plugins for multimodal productivity and Qwen-Live-Harness for building responsive, real-time multimodal agents. These tools enable integration into production workflows for applications such as video editing, long-form translation, and music-conditioned generation.