Alibaba's Qwen team has released Qwen3.8-Omni-Flash, its first omni-modal model designed around agentic capabilities for audio-video understanding and tool use. The model accepts text, images, audio, and video inputs to return text outputs, featuring a 1M-token context window and native function calling.

  • Built on the Qwen3.8-Flash-Next architecture, it offers a 1M-token context with 991K max input and 131K max output tokens.
  • It employs an agentic perception workflow for long videos, raising OmniVideoBench accuracy from 63.4 to 67.8 while reducing token usage by 45.7%.
  • The model shows significant performance gains over Qwen3.5-Omni-Plus, with average scores improving more than 25% across 29 evaluations.
  • It is available via API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio, with pricing at $0.15 per 1M input tokens and $0.47 per 1M output tokens.
  • The team also open-sourced Qwen-MM-Plugins under Apache-2.0 to support multimodal agent harnesses.

The release provides a hosted solution for complex agentic workflows involving long-form audio and video analysis, with reported cost reductions of over 93% for audio-visual inputs compared to previous models.